Package up every result this plugin's tuning/validation research has produced — the layer-ablation study, the prompt-tuning coordinate-ascent passes and the dispatch-floor-awareness/held-out-task work, the averaged-vs-additive `token_ceiling_formula.py` pivot, and the real-work-signal validation experiments — into one condensed, research-paper-style EXECUTIVE report with real charts, built entirely from numbers already recorded in this repo's own dated results files (never invented or rounded up). Publishes a self-contained HTML report (loads the `dataviz` and `artifact-design` skills first) with an abstract, a key-findings table, a handful of figures, limitations stated as prominently as wins, and a reproducibility appendix pointing at the companion skills that can re-run each experiment. Use when someone says "write up all the tuning results", "executive summary of the research", "package the findings into a report", "research report with charts", or "summarize everything we've found so far for leadership".
66
83%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
This plugin's research trail is large and scattered across many dated
files (16+ results files under eval/tuning/results/, two under
eval/ablation/results/, two DESIGN.md narratives, and the source
modules those files report on) — accurate, but not something an executive
reads end to end. This skill's whole job is the compression: read
everything, extract only what's load-bearing, and produce one short,
chart-backed report that states real findings — including the rejected
and null ones — without re-litigating the full narrative.
This is a synthesis skill, not a research skill. It does not run new experiments, dispatch sub-agents, or generate new data — it reports on data that already exists in this repo. If a finding isn't traceable to a specific already-committed file, it doesn't belong in the report.
dataviz before writing a single chart. Its six-check color
validator, form heuristic, and mark specs are not optional polish —
run the palette validator before shipping any chart this skill
produces.artifact-design before writing the HTML file. Calibrate how
much design investment an executive research report warrants (this is
a real deliverable with an audience, not a throwaway) before writing
markup.../../eval/ablation/DESIGN.md +
everything under ../../eval/ablation/results/../../eval/tuning/DESIGN.md +
everything under ../../eval/tuning/results/ (glob it — new files
are added faster than this list can be hand-maintained)../../eval/token_ceiling_formula.py,
../../eval/tuning/knobs.py,
../../eval/tuning/weight_optimizer.py,
../../eval/tuning/overfitting_guard.py
— the source-of-truth constants/verdicts behind the write-ups
(CALIBRATION_STATUS, ADDITIVE_CALIBRATION_STATUS, HOLDOUT_TASKS,
each signal's tested/untested status){name, date, question asked, headline metric before → after, verdict (adopted / rejected / promising-not-proven / structural finding), source file}. This ledger is the report's spine — every
sentence in the sections below should trace back to a row in it. Do
not proceed to writing charts or prose from memory of what a file
"probably said" — pull the exact numbers.eval/ablation/)dispatch_floor_awareness's climb from
level 0 to 3 against real actuals, level 4 and 5's rejection, and the
single-draw-noise correction (0.333 → the true 0.167) that reset how
the whole program measures anything after it.token_ceiling_formula.py pivot — from free-hand integers to
rated signals; the averaged model's PROVEN capacity ceiling (≤16.7%,
provable before training, confirmed by a gradient-checked run); the
additive model's structural fix (94.4%) and the single-scalar
overfitting check that caught it; ADDITIVE_CALIBRATION_STATUS's
honest UNVALIDATED label.validation_loop_iterations
(tested, rejected: dilutes), context_ingestion_volume (tested,
rejected: dilutes, confirmed on genuinely blind data),
investigative_uncertainty (promising, not yet proven — the one
candidate whose combined-sum correlation improved), and the
contamination catch itself (self-authored "blind" draws in a context
already holding the answer produced a fabricated-looking r=0.989,
caught and discarded before it reached a weight).{finding, evidence, verdict} —
scannable in under a minute. Include structural findings (the
capacity-ceiling proof), adopted changes (the additive formula), AND
rejected candidates (both failed signals, both rejected knob levels)
with equal visual weight — a rejected/negative finding is not a
lesser row.ADDITIVE_TOTAL_SPAN's UNVALIDATED status; opus
and haiku tiers' calibration is placeholder-only (scaled by floor
ratio, not independently measured); the self-authored-draw
contamination risk this program already found once and is now
actively guarding against, not something to claim is fully solved.model-right-sizer-holdout-tuning; a second held-out-task check for
investigative_uncertainty via model-right-sizer-signal-validation;
etc.) — not vague "continue researching" language.model-right-sizer-layer-ablation,
model-right-sizer-prompt-tuning, model-right-sizer-holdout-tuning,
model-right-sizer-signal-validation) and, in one line each, what
it re-runs and what ground truth it needs.dataviz's form heuristic for each (a headline number some of
these deserve a stat tile, not a chart) and run the palette validator
before finalizing.results/*.md files are the place for full derivations; this report
links to them (href to a repo path, or name the exact filename) — it
does not restate them.artifact-design),
with a <title>, an accurate description, and a stable favicon. If
the audience needs a static file instead of a link (e.g. for an email
attachment or a printed leave-behind), offer to also produce a PDF or
DOCX version via the pdf/docx skills from the same findings
ledger — but the HTML artifact is the default, since it's the only
format that renders the figures as real, inspectable charts rather
than flattened images.UNVALIDATED calibration status get the same prominence as an
adopted change. This report exists to inform a decision, not to make
the research program look finished.model-right-sizer-layer-ablation,
model-right-sizer-prompt-tuning,
model-right-sizer-holdout-tuning,
model-right-sizer-signal-validation
— the four companion skills that can re-run a piece of this research;
this skill's reproducibility appendix points to all four.../../eval/tuning/DESIGN.md and
../../eval/ablation/DESIGN.md — the
two narrative logs this report condenses.f539a8b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.