CtrlK
BlogDocsLog inGet started
Tessl Logo

panel-review-loop

Iteratively review and improve a Fallow user-facing surface across representative real-world projects until the panel has no blocking concerns.

56

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/panel-review-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplary lean process skill: a tight numbered loop with an explicit stop condition and stability constraints, and no wasted tokens. Its only weakness is that several steps (selecting projects, capturing output, comparing behavior) are stated as goals rather than executable guidance.

Suggestions

Add one concrete example per vague step — e.g., the command used to capture output from a fixture, or how 'compare behavior' is performed (diff, saved transcripts, panel verdicts) — to lift actionability.

Add a short branch for the failure case: when the panel blocks, explicitly tie the fix to the consensus concern and re-run only the affected projects, making the error-recovery loop explicit.

DimensionReasoningScore

Conciseness

The body is ~20 lines with no concept explanations, no padding, and no over-teaching; every line is an instruction or constraint ('Keep the corpus stable across iterations. Preserve output contracts...'), matching the score-5 anchor 'Lean and efficient; assumes Claude's competence; every token earns its place'; there is nothing to trim down to a 4.

5 / 5

Actionability

Two fully concrete commands are given ('npm --prefix benchmarks run download-fixtures', 'Run `panel-review`'), but steps like 'Capture actual output for each project', 'Select representative projects', and 'Implement the smallest coherent improvement' provide no command, example, or criteria for execution, matching the score-3 anchor 'Some concrete guidance but incomplete; missing key details'; it is above a 2 because the setup and review steps are executable as written.

3 / 5

Workflow Clarity

The seven steps are clearly sequenced with an explicit feedback loop ('Run panel-review... Implement the smallest coherent improvement... Re-run the same corpus and compare behavior') and an explicit stop condition ('Stop when the panel has no blocks'), matching the score-4 anchor; it falls short of 5 because error recovery is only implied — there is no explicit instruction for what to do when the panel blocks or how to compare behavior.

4 / 5

Progressive Disclosure

The skill is under 50 lines, single-purpose, needs no external reference files, and is well organized (setup paragraph, numbered procedure, closing constraints), which per the judging guidelines lets a simple skill score 5; the one external pointer ('[benchmark setup](../../../BENCHMARKS.md#comparative-benchmarks)') is a clearly signaled, one-level-deep reference, not buried navigation.

5 / 5

Total

17

/

20

Passed

Description

50%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a clear purpose and end state (panel has no blocking concerns) but omits any 'when to use' trigger guidance and relies on fairly generic verbs (review, improve), leaving it distinguishable mainly by the niche term 'Fallow'. It sits solidly at the midpoint on all dimensions.

Suggestions

Append an explicit trigger clause, e.g. 'Use when iterating on Fallow output quality, when a panel review has raised blocking concerns, or when asked to refine a user-facing surface against real-world projects.'

Add concrete actions and natural synonyms beyond 'review and improve' — e.g. 'capture actual output, run the panel, implement the consensus fix, and re-verify' — to strengthen specificity and trigger-term coverage.

Include the natural vocabulary users would actually say (e.g. 'iterate', 'refine', 'panel feedback', 'real-world fixtures') so the skill triggers for phrasings other than the exact words it currently uses.

DimensionReasoningScore

Specificity

The description names its domain ('a Fallow user-facing surface', 'representative real-world projects') and 1-2 actions ('Iteratively review and improve'), matching the score-3 anchor; it is not a 4 because no further concrete actions (e.g., what the review entails or what outputs change) are listed, and not a 2 because the domain is named with real actions.

3 / 5

Completeness

It clearly answers 'what' ('Iteratively review and improve a Fallow user-facing surface... until the panel has no blocking concerns') but contains no 'Use when...' or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines; not a 2 because the 'what' is clear, not vague, and not a 4 because 'when' is entirely absent rather than just under-specified.

3 / 5

Trigger Term Quality

Relevant keywords exist ('review', 'improve', 'panel', 'real-world projects') but common natural variations users would say ('iterate', 'refine', 'get feedback on', 'polish the UI/CLI output') are missing, matching the score-3 anchor for some-but-incomplete keyword coverage; a 4 would require noticeably broader natural-term coverage.

3 / 5

Distinctiveness Conflict Risk

'Fallow user-facing surface' and 'the panel' give it a project-specific niche, but the generic verbs 'review and improve' could overlap with general code-review or quality-iteration skills, matching the score-3 anchor 'somewhat specific but could still overlap'; not a 4 because without explicit trigger phrases the overlap risk with generic review skills is more than minor.

3 / 5

Total

12

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 suspicious

Warning

Total

15

/

16

Passed

Repository
fallow-rs/fallow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.