CtrlK
BlogDocsLog inGet started
Tessl Logo

model-right-sizer-layer-ablation

Empirically measure what each of model-right-sizer's four research-grounded citation layers (Token Economics, IBPO, BudgetThinker, Speculative Decoding) actually does to its blueprints — instead of trusting the citations alone. Renders layer-ablated variants (any of the 16 layer subsets), runs a fixed six-task benchmark through each variant's Pass A blueprint, and for a scoped subset actually executes the recommended build and scores whether real effort stayed within the predicted budget (wrapping `classify_budget_adherence`). Reports each layer's effect in ISOLATION vs. a zero-layer baseline, and every COMBINATION across the full 16-subset grid, so synergy or redundancy is visible, not assumed away. Read-mostly: writes only a scratch directory and a final report, never `agents/model-right-sizer.md`. Use when someone says "does the Token Economics layer actually change anything", "ablate the research layers", "run the layer-ablation study", or "audit model-right-sizer's citations empirically".

69

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered runbook: correct sequencing, explicit validation and feedback loops at every risky point, cost transparency before the expensive phase, and honest failure-mode guidance grounded in recorded pilot runs rather than generic advice. The two systematic weaknesses are rhetorical over-justification in a few passages and sub-agent dispatch instructions that are described rather than given as a verbatim prompt template.

Suggestions

Tighten the justificatory prose in the Step 2 stall-recovery paragraph and the Step 3 'never fabricate a token count' passage to imperative statements — the rules are valuable but the multi-sentence rationales could be roughly halved without losing the pilot-run citations.

Add a verbatim sub-agent prompt template for one (variant, task) cell in Step 2 (the way model-right-sizer-dryrun instructs it), so the 96 dispatches are copy-paste consistent rather than assembled from a description.

Consider moving the Step 2 stall-diagnosis/recovery procedure and the Step 3 dispatch-overhead normalization guidance into a linked reference file, keeping SKILL.md to the runbook proper.

DimensionReasoningScore

Conciseness

Nearly all of the body is project-specific knowledge Claude cannot already know (pilot-run artifacts, the runtime's 20-dispatch concurrency cap, the ~5-20x whole-dispatch overhead measurement artifact, scope cost math), and the two heredoc scripts are lean. It is not a 5 because several passages carry justificatory padding that could be trimmed — "This is load-bearing, not a style preference", "visible, not assumed away", and the multi-sentence rationales after "Never fabricate a token count" and "a made-up accuracy figure is worse than a visibly incomplete one, because it's indistinguishable from a real one in the report". It stays above anchor 3 because none of it explains concepts Claude already knows — every long passage encodes a pilot-learned failure mode.

4 / 5

Actionability

Two copy-paste-ready executable blocks (the variant-generation python heredoc and the metrics-computation python heredoc), an exact validation command (`uv run --no-project --with jsonschema scripts/validate_blueprint.py -`), exact output paths (`$SCRATCH/results/composition/<variant-name>/<task-id>.json`), and a per-row recipe for the accuracy sweep. It misses anchor 5 only because the sub-agent dispatch instructions in Step 2 are described precisely but never given as a verbatim prompt template, and `SCRATCH=<a scratch directory — never a path inside the repo checkout>` is a placeholder rather than a concrete default — minor gaps, not the missing-steps pattern of anchor 3.

4 / 5

Workflow Clarity

Five clearly sequenced steps with explicit validation and feedback loops throughout: a pre-flight scope confirmation with computed cost counts before the expensive phase; schema validation of every blueprint with a re-emit-once retry loop ("if it doesn't validate, ask that one sub-agent to re-emit once, quoting the validator's error"); a byte-identical check of the all-four variant delegated to `tests/model_right_sizer/test_ablation_layers.py`; a stall-diagnosis and recompute-from-filesystem recovery procedure for batch dispatches; and an explicit "don't infer 'still running' purely from the absence of a failure" checkpoint. This matches the anchor-5 pattern of validate → fix → retry with recovery loops for a batch operation.

5 / 5

Progressive Disclosure

No bundle files exist under references/, scripts/, or assets/ — the body is the whole skill, with all methodology pushed one level out to clearly signaled repo files (`../../eval/ablation/DESIGN.md` — "read that file first; this SKILL.md is the runbook, not the methodology essay" — plus layers.py, benchmark_tasks.json, metrics.py, and a Related section). The body self-describes as a runbook and appropriately defers detail rather than inlining it. It misses anchor 5 because some long inline digressions (the ~40-line stall-recovery procedure and the dispatch-overhead normalization discussion) read as reference material that could be split into their own linked file, and none of the referenced external files are part of this bundle, so a reader outside the parent repo cannot follow them.

4 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concretely what the skill measures and how (16 variants, six-task benchmark, scoped real-build accuracy scoring), discloses its write footprint ("Read-mostly: writes only a scratch directory and a final report"), and gives four natural trigger phrases. The only soft spot is that trigger coverage, while good, is not exhaustive of synonyms a user might spontaneously say.

DimensionReasoningScore

Specificity

The description enumerates the full concrete pipeline: "Renders layer-ablated variants (any of the 16 layer subsets), runs a fixed six-task benchmark through each variant's Pass A blueprint", "actually executes the recommended build and scores whether real effort stayed within the predicted budget (wrapping `classify_budget_adherence`)", and "Reports each layer's effect in ISOLATION vs. a zero-layer baseline, and every COMBINATION across the full 16-subset grid". Coverage spans render, benchmark, execute, score, and report with no gaps — a clear anchor-5 match, not anchor 4's 'minor gaps in coverage'.

5 / 5

Completeness

Both halves are explicit: the 'what' is the measured ablation study described in operational detail, and the 'when' is a literal "Use when someone says ..." clause with four concrete trigger phrases. This matches the anchor-5 example structure exactly (what + 'Use when ...' with concrete triggers), and is clearly above anchor 4's 'when could be more explicit'.

5 / 5

Trigger Term Quality

Four natural quoted trigger phrasings are given ("does the Token Economics layer actually change anything", "ablate the research layers", "run the layer-ablation study", "audit model-right-sizer's citations empirically") — genuinely what a user would say. Not anchor 5 because coverage is strong but not exhaustive: a few plausible natural variants (e.g. "which citation layers matter", "is the BudgetThinker layer real") are absent; not anchor 3 because the quoted phrases already capture the common phrasings users would use for this niche.

4 / 5

Distinctiveness Conflict Risk

It occupies a clear niche — an empirical ablation study of one named agent's (model-right-sizer's) four named citation layers — with distinct triggers unlikely to fire for any other skill. The named layers (Token Economics, IBPO, BudgetThinker, Speculative Decoding) and the wrapping of `classify_budget_adherence` make confusion with a generic benchmarking or sizing skill essentially impossible.

5 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 12 suspicious

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

13

/

16

Passed

Repository
Cloudzero/cloudzero-claude-marketplace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.