CtrlK
BlogDocsLog inGet started
Tessl Logo

perf-loop

Iteratively optimize Fallow performance with stable benchmarks, before-and-after evidence, and correctness gates.

61

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/perf-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplarily lean, well-sequenced optimization loop with genuine validation gates (benchmark sensitivity proof, stop rules, reproducibility and contract checks). The main weakness is actionability: several steps leave operational specifics (profiling tool, statistical thresholds, what counts as a known slowdown) unstated.

Suggestions

Name a concrete profiling command or method for the 'Profile the hot path before editing' step so it is executable rather than directional.

Give quantitative anchors for 'statistically useful baseline' and 'a known slowdown must move the result' (e.g. minimum rounds, a variance threshold, or a percentage change to detect).

Clarify what 'Run review' refers to (which command or skill) so the final step is actionable outside the originating repo.

DimensionReasoningScore

Conciseness

The 23-line body is lean with zero padding: every line is a directive ('Profile the hot path before editing', 'Do not report performance gains from debug builds or incomparable fixtures') and nothing explains concepts Claude already knows.

5 / 5

Actionability

Directives are specific in intent ('Record a statistically useful baseline', 'a known slowdown must move the result') but lack executable detail — no profiling tool is named, no sample-size or variance threshold is given, and `Run review` is the only concrete command — leaving key operational details to inference.

3 / 5

Workflow Clarity

The numbered 1–8 loop has explicit validation checkpoints and feedback gates: prove benchmark sensitivity before freezing, a minimum-rounds/target stop rule, re-run of the same benchmark plus correctness checks, and keep-only-when-reproducible with a contract-regression guard, plus a closing rule against incomparable measurements.

5 / 5

Progressive Disclosure

The skill is under 50 lines, single-purpose, and needs no external references; the compact numbered list plus closing rule is appropriately organized for a self-contained skill with no bundle files present.

5 / 5

Total

18

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise, third-person, and reasonably specific about the optimization methodology, but it omits any 'when to use' trigger guidance and lacks natural trigger synonyms. It communicates what the skill does effectively while leaving discovery to context alone.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks to speed up, optimize, or profile Fallow performance, or reports that Fallow got slower.'

Broaden natural trigger terms with synonyms users would say: 'speed up', 'make it faster', 'slowdown', 'latency', 'benchmark comparison'.

Consider naming one or two concrete actions as verbs (e.g. 'Profiles hot paths and gates each change on reproducible benchmark wins') to strengthen specificity toward anchor 5.

DimensionReasoningScore

Specificity

Concrete third-person phrasing ('stable benchmarks', 'before-and-after evidence', 'correctness gates') anchored to a named domain ('Fallow performance') lists several specific elements, but it describes one methodology rather than enumerating multiple discrete actions, so it falls short of the comprehensive anchor 5.

4 / 5

Completeness

The 'what' is clear (iterative performance optimization with benchmarks, evidence, and gates), but there is no 'Use when...' clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Relevant natural terms ('optimize', 'performance', 'benchmarks') are present, but common variations users would actually say ('speed up', 'make it faster', 'slow', 'latency', 'profile') and any synonyms or extensions are missing.

3 / 5

Distinctiveness Conflict Risk

Scoped to the Fallow project with a distinctive benchmark-evidence-gates methodology, making it mostly distinct with only minor overlap risk against generic profiling or tuning skills.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fallow-rs/fallow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.