CtrlK
BlogDocsLog inGet started
Tessl Logo

conformance-loop

Iteratively improve Fallow analysis accuracy by comparing it with competing tools and verified source truth across a stable real-world corpus.

56

Quality

62%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/conformance-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a concise, well-sequenced workflow with concrete artifacts (a named adjudications file, a classification taxonomy, and a closing review gate) and appropriate verification checkpoints. Its only weakness is minor: steps for running the tools lack execution specifics and there is no explicit fix-retry feedback loop.

DimensionReasoningScore

Conciseness

The body is a lean numbered list that assumes Claude's competence, and the step-5 rationale ('A verdict that lives only in a session or in .plans/ is lost...') earns its place as project-specific motivation rather than concept padding, fitting 'Efficient; minor instances of over-explanation that could be trimmed' rather than the perfectly lean score 5 or the padded scores 1-2.

4 / 5

Actionability

It gives concrete, actionable anchors — a specific output path (`tests/conformance/adjudications.json`), a defined classification taxonomy (true/false positive, false negative, model difference), and an explicit command (`review`) — with only minor gaps on how exactly to run Fallow and the comparison tools, matching 'Mostly executable guidance; concrete code or commands with minor gaps'; the absence of code is not penalized for an instruction-only skill.

4 / 5

Workflow Clarity

An explicit 8-step sequence with verification checkpoints is present (manually verify at step 3, 'retain only net improvements' gate at step 7, 'Run review' at step 8), so the destructive/batch cap does not apply; it fits 'Clear sequence with most checkpoints present; minor validation gaps' rather than score 5 because there is no explicit fix-and-revalidate feedback loop.

4 / 5

Progressive Disclosure

The skill is under 50 lines, has no bundle files, and needs no external references; it is organized as a single well-structured numbered list under one heading, so the simple-skills exception applies and it qualifies for 'Clear overview with well-signaled one-level-deep references' / well-organized short-skill top score.

5 / 5

Total

17

/

20

Passed

Description

46%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly states what the skill does and occupies a distinct, tool-specific niche, but it is jargon-heavy and lacks any explicit 'when to use' trigger guidance. Adding a 'Use when...' clause with natural trigger terms would substantially raise its completeness and trigger-term scores.

Suggestions

Add an explicit 'Use when...' clause, e.g. 'Use when improving Fallow analysis accuracy or adjudicating disagreements between Fallow and competing tools', to satisfy the completeness 'when' requirement.

Replace jargon ('verified source truth', 'real-world corpus') with natural trigger phrases users would actually say, such as 'conformance testing', 'compare Fallow results', or 'verify Fallow findings'.

List a couple more concrete actions (e.g. 'classify true/false positives', 'record adjudications') to lift specificity above the minimal 1-2 action level and broaden trigger coverage.

DimensionReasoningScore

Specificity

It names the domain ('Fallow analysis accuracy') and 1-2 concrete actions ('comparing it with competing tools', 'verified source truth'), but the actions are fairly abstract and coverage is not comprehensive, matching the anchor 'Names domain and 1-2 concrete actions, but not comprehensive' rather than the multi-action score 5 or the minimal score 2.

3 / 5

Completeness

The 'what' is clearly stated (iteratively improve Fallow analysis accuracy via comparison against tools and source truth), but there is no 'Use when...' clause or equivalent explicit trigger guidance, so per the rubric cap completeness cannot exceed 3 ('Has a clear what but when is missing or only weakly implied').

3 / 5

Trigger Term Quality

Phrases like 'verified source truth', 'stable real-world corpus', and 'competing tools' are technical jargon a user would rarely say naturally; there is essentially one generic-ish keyword ('comparing') and the natural phrases users would actually utter ('conformance', 'compare Fallow results', 'verify Fallow findings') are missing, placing it at 'One or two generic keywords; missing the natural phrases users say' rather than the fuller coverage of score 3.

2 / 5

Distinctiveness Conflict Risk

The focus on a named tool ('Fallow analysis accuracy') and a specific conformance/comparison activity gives it a clear niche with only minor overlap risk against closely related evaluation skills, fitting 'Mostly distinct; minor overlap risk' better than the score-3 'could still overlap' or the score-5 'clear niche with distinct triggers' (which requires explicit distinct trigger phrases it lacks).

4 / 5

Total

12

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fallow-rs/fallow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.