Content
73%Weight 40%Scale 1-5Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.
The body delivers a rigorous, well-sequenced experimental protocol with genuine validation gates (blind dispatch, CV diagnostic, combined-sum correlation bar, second-task replication) and strong negative-scope boundaries. Its weaknesses are repetition — the anti-contamination rule and the dilution lesson are each stated three times — and the inline incident narrative that duplicates a linked write-up, which together cost conciseness.
Suggestions
State the anti-contamination rule once (in step 2) and reduce the 'Before starting' incident section to a 3-4 line summary plus the existing link to 2026-08-22-second-signal-experiment-genuinely-blind.md — the full account already lives there.
Add a short copy-paste dispatch prompt template (the exact spec block handed to each sub-agent) so the most fragile step — constructing a genuinely blind draw — is executable rather than descriptive.
Name the concrete mechanism for computing the Pearson correlations (script path, function call, or library invocation), mirroring how the test-file and registry paths are already given exactly.
| Dimension | Reasoning | Score |
|---|---|---|
Conciseness | Quotes: the contamination incident is narrated at length ('A first attempt... generated all three "blind" draws inline... it's transcribing the answer key and then reporting that it correlates with the answer') and then restated in step 2 ('Do not generate the draws yourself inline, even carefully, even if you believe you can reason about it "as if blind"') and again in the NOT-do section ('It does not accept an inline, self-authored "blind" draw under any framing — see the incident above'). The dilution lesson likewise appears in the intro, step 4, and step 5. The content is project-specific, not concepts Claude already knows, so it is not a 2, but the same critical rules are stated three times where once plus a pointer to the linked incident write-up would carry the same force. | 3 / 5 |
Actionability | Quotes: 'Dispatch 3 or more separate Task/Agent calls, each given only: the full definition + 0.0/1.0 anchors for every signal being rated', 'average each signal's value across the draws. Compute... the coefficient of variation (stdev / mean across the draws)', 'the candidate signal alone; the existing signal sum... plus the candidate at equal weight', and 'Add regression tests mirroring tests/model_right_sizer/test_token_ceiling_formula.py's existing test_*_dilutes_the_existing_signal_correlation pattern'. This is concrete, executable guidance with exact file paths, thresholds, and naming patterns — as an instruction-only skill, absence of code is not penalized per the scoring notes. It is not a 5 because there is no copy-paste-ready dispatch prompt template (the single most fragile step) and no command showing how the correlation is actually computed (script call or function name). | 4 / 5 |
Workflow Clarity | Quotes: an 8-step numbered sequence from 'Pick the candidate signal and a held-out task' through 'A weight change is a proposed diff for a human to review', with explicit validation checkpoints throughout: CV as 'the noise-magnitude diagnostic, not a pass/fail gate', the combined-sum correlation bar ('Judge a candidate on the combined-sum delta, not the standalone figure'), the replication gate ('Only propose... once the SAME candidate clears the correlation bar on a second, different held-out task'), and error-recovery paths ('if no fresh task exists yet, say so explicitly and treat the result as weaker evidence'; 'If the runtime has no sub-agent dispatch available at all, stop and say so'). This matches the anchor-5 shape: clear sequence, explicit validation, feedback loops for failure modes. | 5 / 5 |
Progressive Disclosure | The body is organized into clear sections (incident, 'What to do' steps, 'What this does NOT do', 'Related') and the Related section lists well-signaled one-level-deep references with markdown links, e.g. '[`../../eval/token_ceiling_formula.py`]... — the module under test', '[`../../eval/tuning/overfitting_guard.py`]... — HOLDOUT_TASKS'. No nested-reference chains exist, and no bundle files are provided to score against. It is not a 5 because the ~30-line incident narrative largely duplicates the linked results write-up ('2026-08-22-second-signal-experiment-genuinely-blind.md') and could live in a reference file with a two-line summary in the body. | 4 / 5 |
Total | 16 / 20 Passed |