CtrlK
BlogDocsLog inGet started
Tessl Logo

test-provenance-guard

Detects tests that pass by construction — tests that define a private copy of the function under test instead of importing the production module — and self-heals by extracting the inline logic to an exported function, updating production callers, and rewriting the test to import the export. Two checks: (1) static — the test file must import the SUT and must not shadow its exported names; (2) mutation — blanking the production function body re-runs the test and expects failure. Runs autonomously inside autonomous-workflow Phase 4 and as a slash command for human-driven PR review. Use when adding new tests for existing or refactored code, when CI is green but you are unsure whether the tests actually exercise production, or when reviewing a PR for tests-by-construction. Triggers on "test provenance", "tests by construction", "verify tests cover real code", "tests duplicate logic", "mutation sanity check", "are these tests fake", "/test-provenance-guard".

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-conceived thin-index skill body with clear phasing, gates, and an output contract, undermined by a broken bundle: the three rule files that carry the actual detection and self-heal procedures are missing, so the skill cannot be executed as written. Redundant restatement of the confidence gate is the main token-efficiency drag.

Suggestions

Add the missing rules/static-check.md, rules/mutation-check.md, and rules/self-heal.md files (or inline their procedures) — the Required Reading by Phase table and the Phase-gate workflow currently point to nonexistent files.

State the confidence ≥ 90 % gate once in the Inputs section and reference it elsewhere ('see Inputs: --fix') instead of restating the threshold in the workflow table, decision flow, output contract, and Definition of Done.

Replace or supplement the pseudocode Quick Decision Flow with the concrete detection heuristic for at least the static check (e.g., what counts as a shadowed export) so the cheapest check is executable even without the rule files.

DimensionReasoningScore

Conciseness

The body is largely lean and assumes competence ('No test files? Exit 0.'), but the confidence gate (≥ 90 %) is restated in the Inputs table, the paragraph below it, the Workflow table, the Quick Decision Flow, the Output Contract, and the Definition of Done — roughly six repetitions of the same fact. Anchor 5 requires every token to earn its place; this repetition could be consolidated into one canonical statement.

4 / 5

Actionability

Some concrete guidance exists (argument table with defaults, `git diff --name-only <base>...HEAD`, `git restore <sut-file>`, structured report template), but the actual executable procedures — how static_check detects shadowed exports, how the mutation is performed and restored, how the extraction is planned — all live in `rules/static-check.md`, `rules/mutation-check.md`, and `rules/self-heal.md`, none of which exist in the bundle. The Quick Decision Flow is explicitly pseudocode. Key execution details are therefore missing, matching the 'incomplete / missing key details' anchor rather than the 'minor gaps' one.

3 / 5

Workflow Clarity

The three phases are clearly sequenced with explicit gates ('Phase 2 runs only when Phase 1 passes', mutation check re-run after heal, '<TEST_CMD> re-run is green') and a Definition of Done checklist — feedback loops are present for this mutating operation. Not a 5: the per-phase operational detail is delegated to rule files that are absent from the bundle, and the pseudocode flow's `Skill("confidence", ...)` call lacks a concrete invocation contract.

4 / 5

Progressive Disclosure

The thin-index design is correct in principle — 'This SKILL.md is a thin index... Detailed procedures live in rules/*.md and load on demand', one level deep, clearly signaled — but scored against the actual bundle structure, all three referenced rule files (static-check.md, mutation-check.md, self-heal.md) do not exist; only the two references/*.md files are present. The 'Required Reading by Phase' table sends the reader to dead links for every phase, leaving the skill's core procedure unreachable. Structure is more than 'minimal' but the broken reference graph pulls it below the midpoint.

2 / 5

Total

13

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A model description: concrete capabilities, both checks spelled out, explicit 'Use when' guidance with natural trigger phrases, third-person voice, and a distinct niche with minimal conflict risk. Dense without padding — every clause carries information.

DimensionReasoningScore

Specificity

Multiple concrete actions are enumerated with precision: 'static — the test file must import the SUT and must not shadow its exported names', 'mutation — blanking the production function body re-runs the test and expects failure', and self-heal steps 'extracting the inline logic to an exported function, updating production callers, and rewriting the test to import the export'. Coverage is comprehensive across detection and remediation; not merely naming the domain.

5 / 5

Completeness

What: 'Detects tests that pass by construction... and self-heals by extracting the inline logic...'. When: an explicit 'Use when adding new tests for existing or refactored code, when CI is green but you are unsure whether the tests actually exercise production, or when reviewing a PR' clause with concrete trigger phrases. Both halves are explicit and concrete.

5 / 5

Trigger Term Quality

Natural phrases a user would actually say are explicitly listed: 'are these tests fake', 'tests duplicate logic', 'verify tests cover real code', 'mutation sanity check', plus 'test provenance' and the slash form '/test-provenance-guard'. Both expert and casual phrasings are covered.

5 / 5

Distinctiveness Conflict Risk

The niche (tests-by-construction detection with mutation-based verification) is highly specific and distinct from generic test-quality or code-review skills; triggers like 'tests by construction' and 'are these tests fake' are unlikely to fire for the wrong skill.

5 / 5

Total

20

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 7 missing, 4 suspicious

Warning

Total

13

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.