CtrlK
BlogDocsLog inGet started
Tessl Logo

conformance-loop

Iteratively improve Fallow analysis accuracy by comparing it with competing tools and verified source truth across a stable real-world corpus.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/conformance-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplary lean process skill: a clearly sequenced loop with a record-before-fix gate and a corpus-level validation step, and no wasted tokens. Its only weaknesses are a few steps that assume context the reader may not have (corpus definition, the format of the adjudications record, what "Run `review`" entails).

DimensionReasoningScore

Conciseness

A lean 8-step numbered process with zero padding and no explanation of concepts Claude already knows; the one explanatory sentence in step 5 (why the verdict must be recorded) earns its place by preventing the rule from being skipped. It is not a 4 because there is no trimmable content that would not lose information.

5 / 5

Actionability

For an instruction-only skill the guidance is mostly executable: an exact output path (tests/conformance/adjudications.json), an explicit classification vocabulary (true/false positive, false negative, model difference), and a required regression fixture. Not a 5 because steps like "Define the capability and stable project corpus" and "Run `review`" give direction without the specific details needed to execute them.

4 / 5

Workflow Clarity

The 8-step sequence is clearly ordered with real checkpoints: step 5 gates the fix on a recorded adjudication, and step 7's "re-run the full corpus and retain only net improvements" is an explicit validation with rollback for the batch change. Not a 5 because recovery behavior when the fix regresses is only implicit and step 8 is a bare invocation with no criteria.

4 / 5

Progressive Disclosure

The skill is under 50 lines, single-purpose, needs no external reference files (none exist in the bundle), and is well-organized as a numbered list — the rubric's simple-skill exception applies directly. The only path mentioned (tests/conformance/adjudications.json) is an output location, not a nested reference.

5 / 5

Total

18

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states a clear and fairly specific capability tied to a named tool, but it has no 'when to use' trigger guidance and lacks the natural trigger vocabulary that would let a user (or Claude) reliably select this skill. It reads as a competent one-sentence summary rather than an effective routing description.

Suggestions

Add an explicit trigger clause, e.g. "Use when adjudicating disagreements between Fallow and other static-analysis tools, or when building a conformance/benchmark corpus for analyzer accuracy."

Include natural synonyms users would say — "conformance testing", "benchmark", "false positives/false negatives", "ground truth" — to improve trigger matching.

Split the single composite action into 2-3 distinct concrete actions (e.g. "compare outputs across tools, verify each disagreement against source, and record adjudications") to raise specificity.

DimensionReasoningScore

Specificity

Names the domain (Fallow analysis accuracy) and a concrete method ("comparing it with competing tools and verified source truth across a stable real-world corpus"), but this is one composite action rather than a list of several specific actions, so it is not comprehensive. It is above a 2 because the actions described are specific rather than generic.

3 / 5

Completeness

The 'what' is clear (improve Fallow analysis accuracy by comparison against competing tools and source truth), but there is no "Use when..." clause or any equivalent explicit trigger guidance; the judging guidelines state a missing 'when' caps completeness at 3.

3 / 5

Trigger Term Quality

Contains relevant keywords ("Fallow", "competing tools", "accuracy", "corpus"), but misses the natural phrases a user would actually say when needing this skill — e.g. "conformance", "benchmark", "false positives", "ground truth". Anchor 3 ("some relevant keywords but missing common variations or synonyms") is the best fit.

3 / 5

Distinctiveness Conflict Risk

Naming the specific tool "Fallow" carves a clear niche with minimal conflict risk, but the generic framing ("iteratively improve accuracy") leaves minor overlap with general benchmarking or testing skills, placing it between anchors 3 and 5.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fallow-rs/fallow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.