CtrlK
BlogDocsLog inGet started
Tessl Logo

model-right-sizer-research-report

Package up every result this plugin's tuning/validation research has produced — the layer-ablation study, the prompt-tuning coordinate-ascent passes and the dispatch-floor-awareness/held-out-task work, the averaged-vs-additive `token_ceiling_formula.py` pivot, and the real-work-signal validation experiments — into one condensed, research-paper-style EXECUTIVE report with real charts, built entirely from numbers already recorded in this repo's own dated results files (never invented or rounded up). Publishes a self-contained HTML report (loads the `dataviz` and `artifact-design` skills first) with an abstract, a key-findings table, a handful of figures, limitations stated as prominently as wins, and a reproducibility appendix pointing at the companion skills that can re-run each experiment. Use when someone says "write up all the tuning results", "executive summary of the research", "package the findings into a report", "research report with charts", or "summarize everything we've found so far for leadership".

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable content with an excellent workflow: exact schemas, a fully specified report shape, real checkpoints, and disciplined one-level-deep references to the underlying sources. The main weakness is redundancy — the same non-goals and the equal-prominence-for-negative-findings principle are each stated multiple times, inflating token cost.

Suggestions

Consolidate the redundancy between the intro ('This is a synthesis skill... does not run new experiments') and 'What this does NOT do' into a single statement, and state the 'rejected findings get equal prominence' principle once instead of three times.

Trim step 2's inline numeric detail (e.g. the 0.333→0.167 and 94.4% specifics) to experiment names plus questions, since the ledger step already mandates pulling exact numbers from source files — this also reduces staleness risk when results files change.

Add one line of local guidance for palette-validator failures (e.g. 'fix the chart per dataviz's validator output and re-run before finalizing') so the verification loop does not depend entirely on the dataviz skill.

DimensionReasoningScore

Conciseness

Mostly efficient, but the same points are argued repeatedly: the intro's 'It does not run new experiments, dispatch sub-agents, or generate any new data' is restated nearly verbatim in 'What this does NOT do', and the 'rejected findings get equal prominence' principle is explained three separate times ('limitations stated as prominently as wins', 'a rejected/negative finding is not a lesser row', 'It does not average away or soften a rejected finding'). These could be tightened without losing content.

3 / 5

Actionability

For an instruction-only skill this is copy-paste-level specific: an exact ledger schema '{name, date, question asked, headline metric before → after, verdict (adopted / rejected / promising-not-proven / structural finding), source file}', a section-by-section report shape with sentence/row budgets, named source file paths, and a concrete figure list anchored to real numbers ('0.333 was later corrected to the true 0.167', '~16.7% vs. 94.4%'). Flexibility ('adjust to whatever the ledger actually supports') is explicitly justified.

5 / 5

Workflow Clarity

A clear sequence (load dataviz/artifact-design → read the full source inventory → build the ledger → group into narrative arc → write the report → figures → publish) with real checkpoints: 'Do not proceed to writing charts or prose from memory', 'run the palette validator before finalizing', and a recovery path for missing data ('say so in the figure's caption rather than interpolate invented points'). Not 5 because what to do when the palette validator fails is deferred entirely to the dataviz skill with no local guidance.

4 / 5

Progressive Disclosure

No bundle files exist, and the body's external references (eval DESIGN.md/results files, four sibling companion skills) are one level deep, clearly signaled as links, and the body explicitly refuses to restate them ('link to the source file for anyone who wants the proof'). Minor gap: step 2's per-experiment narrative descriptions inline substantial numeric detail that duplicates what the mandated ledger step already pulls from the source files.

4 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit what-and-when with natural quoted triggers, third person throughout, and a distinct niche. Its only weakness is verbosity — heavy enumeration of internal experiment names inflates length without improving clarity or trigger coverage.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions in third person ('Package up every result... into one condensed, research-paper-style EXECUTIVE report', 'Publishes a self-contained HTML report... with an abstract, a key-findings table, a handful of figures, limitations... and a reproducibility appendix'), matching the comprehensive-coverage anchor. It falls short of 5 only because the dense enumerations of internal experiment names ('coordinate-ascent passes', 'averaged-vs-additive token_ceiling_formula.py pivot') pad the length without adding action clarity.

4 / 5

Completeness

It explicitly answers both: what ('Package up every result... into one condensed, research-paper-style EXECUTIVE report with real charts, built entirely from numbers already recorded in this repo's own dated results files') and when ('Use when someone says "write up all the tuning results"...'), with concrete trigger phrases — a clear match to the top anchor.

5 / 5

Trigger Term Quality

The quoted triggers are natural phrases a user would actually say: "write up all the tuning results", "executive summary of the research", "package the findings into a report", "research report with charts", "summarize everything we've found so far for leadership". Not 5 because common variations (e.g. 'report the findings', 'write-up', 'brief leadership on the results') are missing and much of the remaining keyword mass is internal jargon rather than user-said terms.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche — synthesizing already-committed research results into an executive report, explicitly 'never invented or rounded up' — that is cleanly separated from the companion experiment-running skills, and its triggers ('executive summary of the research', 'research report with charts') would not naturally fire for those. Minimal conflict risk.

5 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 12 suspicious

Warning

Total

14

/

16

Passed

Repository
Cloudzero/cloudzero-claude-marketplace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.