CtrlK
BlogDocsLog inGet started
Tessl Logo

binary-loop

Iteratively reduce Fallow binary size using cargo-bloat and release-build measurements while preserving features, performance, and compatibility.

61

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/binary-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally lean, well-sequenced iterative workflow with an explicit keep/discard checkpoint and a scope guardrail. The main weakness is actionability: the measurement and build steps name tools but provide no concrete commands, and the failure path (revert on regression) is left implicit.

Suggestions

Add the concrete measurement commands, e.g. `cargo build --release && cargo bloat --release --crates` and how to record before/after binary sizes (e.g., `stat`/`ls -l` on the artifact).

Make the discard path explicit: "If size regresses or verification fails, revert the change before selecting the next candidate."

Specify what "verification passes" means operationally (test suite command, target-platform checks) so step 5's gate is executable rather than aspirational.

DimensionReasoningScore

Conciseness

The body is ~15 lines with zero padding: every step ("Measure the release binary and capture cargo-bloat evidence", "Rebuild with identical release settings") earns its place and no concepts Claude already knows are explained, matching anchor 5 ("Lean and efficient; assumes Claude's competence; every token earns its place"). It is not anchor 4 because there are no instances of over-explanation to trim.

5 / 5

Actionability

The workflow is conceptually concrete — naming cargo-bloat, release settings, and candidate categories ("dependency, monomorphization, feature, or codegen contributor") — but no executable commands are given (e.g., the actual `cargo bloat --release` invocation, the build command, or how to measure and compare binary size), matching anchor 3 ("Some concrete guidance but incomplete; missing key details"). It is not anchor 4 because a practitioner could not run the measurement step as written without filling in commands themselves.

3 / 5

Workflow Clarity

The seven steps form a clear, well-sequenced loop with an explicit validation gate ("Keep the change only when size improves and verification passes") and a termination criterion ("Repeat until the target is met or remaining candidates have poor tradeoffs"), matching anchor 4 ("Clear sequence with most checkpoints present; minor validation gaps"). It falls short of anchor 5 because the error-recovery path is implicit — there is no explicit 'if verification fails, revert the change and try the next candidate' instruction, and how verification is performed is unspecified.

4 / 5

Progressive Disclosure

This is a simple skill under 50 lines with no bundle files (no references/, scripts/, or assets/ directories exist), and the body is a single well-organized section whose content is appropriately complete in-line; per the rubric's simple-skill guideline this matches anchor 5. The only cross-reference (`review`) is to another skill/command, not a nested bundle file, so there are no buried or multi-level references.

5 / 5

Total

17

/

20

Passed

Description

65%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, concrete, third-person description with excellent distinctiveness and good trigger terms, but it completely lacks a 'when to use' clause, which caps completeness and leaves the trigger context implicit. Adding an explicit usage trigger would lift it from good to excellent.

Suggestions

Add an explicit trigger clause, e.g. "Use when the Fallow release binary is too large, when asked to shrink or optimize binary size, or before cutting a release."

Include common user phrasings as trigger synonyms — "shrink the binary", "strip the binary", "binary bloat" — to broaden natural-term coverage.

Optionally enumerate the concrete actions already implied by the body (measure with cargo-bloat, select a contributor, make a bounded change, revert if size or verification regresses) to strengthen the 'what' coverage.

DimensionReasoningScore

Specificity

The description names the domain and method concretely — "reduce Fallow binary size using cargo-bloat and release-build measurements" — but describes a single goal with its instrumentation rather than listing several distinct actions, matching anchor 3 ("Names domain and 1-2 concrete actions, but not comprehensive"). It does not reach anchor 4 because there is no enumeration of multiple operations (e.g., measuring, selecting contributors, reverting) that a fuller description would list.

3 / 5

Completeness

The 'what' is clear — iteratively reduce binary size with cargo-bloat and release measurements while preserving features, performance, and compatibility — but there is no 'when' clause at all (no "Use when..." or equivalent trigger guidance), which caps completeness at 3 per the judging guidelines. This matches anchor 3 ("Has a clear 'what' but 'when' is missing or only weakly implied"); it cannot be 4 without any explicit usage-trigger phrasing.

3 / 5

Trigger Term Quality

Strong natural terms are present — "binary size", "reduce", "cargo-bloat", "release-build" — which a user asking to slim a Rust binary would plausibly say, matching anchor 4 ("Good keyword coverage; a few natural terms missing"). It falls short of anchor 5 because common synonyms such as "shrink", "strip", "slim down", or "binary is too big" are absent.

4 / 5

Distinctiveness Conflict Risk

The description carves out a clear niche — Fallow binary size reduction via cargo-bloat — with minimal overlap risk against other skills, matching anchor 5 ("Clear niche with distinct triggers; minimal conflict risk"). Anchor 4 would require some overlap with closely related skills, which the project-specific scoping here avoids.

5 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fallow-rs/fallow
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.