CtrlK
BlogDocsLog inGet started
Tessl Logo

conformance-loop

Iteratively improve Fallow analysis accuracy by comparing it with competing tools and verified source truth across a stable real-world corpus.

62

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/conformance-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplar of lean, imperative skill writing: a tight 8-step loop with concrete artifacts, an explicit adjudication record, and built-in verification and net-improvement gates. The only soft spot is that a few steps ('documented equivalent settings', 'regression fixture') assume project context without pointing to where it lives.

DimensionReasoningScore

Conciseness

The ~20-line body is lean with zero padding and no explanation of concepts Claude already knows; the rationale in step 5 is project-specific behavioral enforcement ('the daily lane then re-reports the same disagreement as unadjudicated forever'), not filler, so every token earns its place.

5 / 5

Actionability

Concrete guidance throughout: a specific artifact path ('tests/conformance/adjudications.json'), an explicit classification taxonomy ('true positive, false positive, false negative, or model difference'), and a named final command ('review'). Minor gaps remain, e.g. 'documented equivalent settings' does not say where or how settings are documented, keeping it below fully-executable.

4 / 5

Workflow Clarity

A clear 8-step sequence with explicit validation checkpoints (step 3 'Manually verify disagreements against source') and a genuine feedback loop (step 7 'Re-run the full corpus and retain only net improvements'), plus a final review gate — matching the top anchor for sequenced validation with error-recovery loops.

5 / 5

Progressive Disclosure

The skill is under 50 lines, single-purpose, with no bundle files present and no external references needed; the single well-organized numbered procedure under one heading fully qualifies under the simple-skill exception.

5 / 5

Total

19

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a specific niche purpose but reads as internal documentation rather than a trigger-facing description. It lacks a 'Use when...' clause and natural trigger phrases, so it will rarely be surfaced by user requests even when the skill is exactly what is needed.

Suggestions

Append an explicit trigger clause, e.g. 'Use when Fallow's findings disagree with another tool, when validating Fallow against ground truth, or when the user mentions conformance, benchmarking, or adjudicating Fallow results.'

Replace internal jargon ('verified source truth', 'stable real-world corpus') with natural synonyms users would say, such as 'ground truth', 'benchmark', 'test corpus', or 'real projects'.

Enumerate the concrete actions more fully in the description (run both tools, adjudicate disagreements as TP/FP/FN, record verdicts, fix and regression-test) to lift specificity above the 1-2 action level.

DimensionReasoningScore

Specificity

Names the domain ('Fallow analysis accuracy') and 1-2 concrete mechanisms ('comparing it with competing tools', 'verified source truth', 'stable real-world corpus'), but does not enumerate multiple specific actions comprehensively as a 4 or 5 would.

3 / 5

Completeness

The 'what' is clear (iteratively improve Fallow analysis accuracy via comparison and verification), but there is no 'Use when...' clause or equivalent trigger guidance, capping completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Contains some relevant domain keywords ('competing tools', 'analysis accuracy', 'corpus') but misses natural phrases a user would say such as 'benchmark', 'conformance', 'evaluate', or 'ground truth', leaning on internal jargon like 'verified source truth'.

3 / 5

Distinctiveness Conflict Risk

The named tool ('Fallow') and conformance-corpus methodology give it a clear niche with minimal conflict risk, but the absence of explicit distinct triggers leaves minor overlap with general benchmarking or code-review skills.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fallow-rs/fallow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.