CtrlK
BlogDocsLog inGet started
Tessl Logo

perf-loop

Iteratively optimize Fallow performance with stable benchmarks, before-and-after evidence, and correctness gates.

62

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/perf-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary lean, well-sequenced performance-optimization loop with real validation gates and no token waste. Its one weakness is actionability: the methodology is sound but operators get no concrete commands or thresholds for baselining, profiling, or judging reproducibility.

Suggestions

Name concrete tooling or commands for the measurement steps, e.g. how to run the stable benchmark, how to profile the hot path, and a concrete minimum round count for 'statistically useful'.

Add an explicit failure branch after step 6: what to do when the improvement is not reproducible or a contract regresses (e.g., 'Revert the change, re-run the correctness checks, and try a different optimization').

Define 'reproducible' operationally (e.g., improvement holds across N re-runs or exceeds a stated threshold) so the keep/discard gate is decidable.

DimensionReasoningScore

Conciseness

Twenty lean lines with zero padding or explanation of concepts Claude already knows; every directive ('Profile the hot path before editing', 'Do not report performance gains from debug builds') earns its place. No neighbor anchor fits better.

5 / 5

Actionability

Guidance is directive and unambiguous ('prove that it can see the problem: a known slowdown must move the result', 'Run `review`'), but key execution details are missing — no commands or concrete methods for recording a 'statistically useful baseline', profiling, or deciding what 'reproducible' means. Not 4 because the only literal executable element is the `review` command.

3 / 5

Workflow Clarity

A clear 8-step sequence with explicit validation gates (step 1 proves the benchmark, step 5 re-runs benchmarks and correctness checks, step 6 gives a keep-only-when-reproducible criterion). Not 5 because there is no explicit error-recovery branch — what to do when correctness regresses or the improvement is not reproducible is implied rather than stated.

4 / 5

Progressive Disclosure

The skill is under 50 lines, needs no external references (none exist in the bundle), and is well organized as a numbered loop with a closing guardrail note, which scores 5 under the simple-skill exception in the scoring notes.

5 / 5

Total

17

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, reasonably distinctive description with good natural trigger terms, but it omits any 'when to use' guidance, which is the main gap. Adding an explicit trigger clause would likely lift completeness and trigger coverage to the top anchors.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks to speed up or optimize Fallow, reports a slowdown, or requests benchmark-backed performance work.'

Include natural synonyms users would actually say — 'faster', 'speed up', 'latency', 'profiling', 'regression' — to broaden trigger term coverage.

State when NOT to use it (e.g., non-Fallow targets or one-off timing checks) to sharpen distinctiveness against generic performance skills.

DimensionReasoningScore

Specificity

Names the domain ('optimize Fallow performance') and several concrete mechanisms — 'stable benchmarks', 'before-and-after evidence', 'correctness gates'. It falls short of 5 because the actions describe an approach rather than a comprehensive list of capabilities, and exceeds 3 because more than 1-2 concrete actions are given.

4 / 5

Completeness

The 'what' is clear (iterative performance optimization with benchmarks, evidence, and gates), but there is no 'Use when...' clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines. It is not 2 because the 'what' is specific rather than vague.

3 / 5

Trigger Term Quality

'optimize', 'performance', and 'benchmarks' are natural phrases a user would say when needing this skill. Not 5 because common variations like 'faster', 'speed up', 'latency', 'profiling', or 'regression' are missing.

4 / 5

Distinctiveness Conflict Risk

Tied to a specific target ('Fallow') with a distinctive methodology (benchmarks + evidence + correctness gates), so it is mostly distinct. Not 5 because generic perf-tuning requests ('make this faster') could still route to it or competing performance skills.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fallow-rs/fallow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.