CtrlK
BlogDocsLog inGet started
Tessl Logo

ark-chainsaw-testing

Run and write Ark Chainsaw tests with mock-llm. Use for running tests, debugging failures, or creating new e2e tests.

88

1.37x
Quality

85%

Does it follow best practices?

Impact

95%

1.37x

Average score across 3 eval scenarios

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, highly actionable skill body: real commands, a concrete test layout, and three well-justified antipatterns with correct/incorrect code. Main gaps are minor — slight prose/comment redundancy, an implicit fix-rerun loop, and bundle references that don't resolve within the skill directory.

DimensionReasoningScore

Conciseness

The body is dense with non-obvious, repo-specific knowledge (kubectl nil-conditions bug, chainsaw assert-vs-watch semantics) and avoids explaining basics Claude knows. Minor trimming is possible — inline comments like '# Bad - polls every few seconds, hammers the API server' duplicate the surrounding prose — so it fits anchor 4 (efficient, minor over-explanation) rather than anchor 5's every-token-earns-its-place.

4 / 5

Actionability

Everything is executable: copy-paste chainsaw commands ('chainsaw test ./tests/query-parameter-ref --fail-fast', '--skip-delete --pause-on-failure'), a concrete directory layout, and complete bad/good YAML pairs for each antipattern. Matches anchor 5 (fully executable, covers common cases); clearly above anchor 4's 'minor gaps'.

5 / 5

Workflow Clarity

Running is sequenced with built-in feedback mechanisms (--fail-fast, then --pause-on-failure/--skip-delete for debugging), and the assert-after-wait pattern is an explicit two-step validation sequence. It falls short of anchor 5 because the run-fail-debug-fix-rerun loop and the test-writing order are implied rather than spelled out; it exceeds anchor 3 since debug flags and separate wait/assert steps act as checkpoints.

4 / 5

Progressive Disclosure

Sections are well organized and external material is signaled one level deep ('Reference tests/CLAUDE.md for comprehensive patterns', 'see [examples.md](examples.md)'). It fits anchor 4 (good structure, minor gaps) rather than 5 because the referenced examples.md/tests/CLAUDE.md are not present in the skill bundle itself, and the three long antipattern subsections (~half the file) could arguably live in a reference file.

4 / 5

Total

17

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete capabilities, third-person voice, and an explicit 'Use for...' clause with three trigger scenarios. The niche domain terms make conflict risk minimal; only minor synonym coverage keeps trigger quality from a perfect score.

DimensionReasoningScore

Specificity

The description names the domain ('Ark Chainsaw tests with mock-llm') and several concrete actions — 'Run and write' tests plus 'running tests, debugging failures, or creating new e2e tests'. It matches anchor 4 (several specific actions, minor gaps) rather than 5, since it omits capabilities like fixing/flaking triage; not anchor 3 because more than 1-2 actions are named across the two clauses.

4 / 5

Completeness

Both parts are explicit: 'Run and write Ark Chainsaw tests with mock-llm' answers what, and 'Use for running tests, debugging failures, or creating new e2e tests' answers when with three concrete trigger scenarios. This matches anchor 5; it exceeds anchor 4 because the 'when' clause is explicit and specific rather than something that 'could be more explicit'.

5 / 5

Trigger Term Quality

Natural phrases users would say are present — 'running tests', 'debugging failures', 'creating new e2e tests', plus domain terms 'Chainsaw' and 'mock-llm'. A few natural variants are missing (e.g., 'end-to-end' spelled out, 'test failures', 'flaky tests'), so it fits anchor 4 rather than anchor 5's comprehensive synonym coverage, and is well above anchor 3's missing-common-variations.

4 / 5

Distinctiveness Conflict Risk

'Ark Chainsaw tests with mock-llm' carves out a clear niche with distinct, repo-specific triggers; no generic phrasing that would collide with other testing skills. Matches anchor 5 (clear niche, minimal conflict risk) and clearly above anchor 4's minor-overlap example.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
mckinsey/agents-at-scale-ark
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.