CtrlK
BlogDocsLog inGet started
Tessl Logo

model-right-sizer-signal-validation

Test whether a candidate real-work signal in `eval/token_ceiling_formula.py` (e.g. `context_ingestion_volume`, `investigative_uncertainty`, or new) deserves a nonzero default weight, via the blind multi-draw rating + correlation methodology this repo's research converged on. Ratings must come from genuinely independent sub-agent dispatches seeing ONLY a forward-looking task spec and the signal definitions — never a context holding the real actuals or this repo's retired write-ups, which would turn "blind rating" into transcribing the answer key. Dispatches 3+ independent draws per held-out task, computes per-signal CV and Pearson correlation (alone, and added to the signal sum — dilution, not weak correlation, is the dominant failure found twice), requiring replication on a SECOND held-out task before proposing a nonzero weight. Use when someone says "test this new signal", "does [signal] deserve a nonzero weight", "re-run the signal validation experiment", or "validate real-work signals against real data".

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

73%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a rigorous, well-sequenced experimental protocol with genuine validation gates (blind dispatch, CV diagnostic, combined-sum correlation bar, second-task replication) and strong negative-scope boundaries. Its weaknesses are repetition — the anti-contamination rule and the dilution lesson are each stated three times — and the inline incident narrative that duplicates a linked write-up, which together cost conciseness.

Suggestions

State the anti-contamination rule once (in step 2) and reduce the 'Before starting' incident section to a 3-4 line summary plus the existing link to 2026-08-22-second-signal-experiment-genuinely-blind.md — the full account already lives there.

Add a short copy-paste dispatch prompt template (the exact spec block handed to each sub-agent) so the most fragile step — constructing a genuinely blind draw — is executable rather than descriptive.

Name the concrete mechanism for computing the Pearson correlations (script path, function call, or library invocation), mirroring how the test-file and registry paths are already given exactly.

DimensionReasoningScore

Conciseness

Quotes: the contamination incident is narrated at length ('A first attempt... generated all three "blind" draws inline... it's transcribing the answer key and then reporting that it correlates with the answer') and then restated in step 2 ('Do not generate the draws yourself inline, even carefully, even if you believe you can reason about it "as if blind"') and again in the NOT-do section ('It does not accept an inline, self-authored "blind" draw under any framing — see the incident above'). The dilution lesson likewise appears in the intro, step 4, and step 5. The content is project-specific, not concepts Claude already knows, so it is not a 2, but the same critical rules are stated three times where once plus a pointer to the linked incident write-up would carry the same force.

3 / 5

Actionability

Quotes: 'Dispatch 3 or more separate Task/Agent calls, each given only: the full definition + 0.0/1.0 anchors for every signal being rated', 'average each signal's value across the draws. Compute... the coefficient of variation (stdev / mean across the draws)', 'the candidate signal alone; the existing signal sum... plus the candidate at equal weight', and 'Add regression tests mirroring tests/model_right_sizer/test_token_ceiling_formula.py's existing test_*_dilutes_the_existing_signal_correlation pattern'. This is concrete, executable guidance with exact file paths, thresholds, and naming patterns — as an instruction-only skill, absence of code is not penalized per the scoring notes. It is not a 5 because there is no copy-paste-ready dispatch prompt template (the single most fragile step) and no command showing how the correlation is actually computed (script call or function name).

4 / 5

Workflow Clarity

Quotes: an 8-step numbered sequence from 'Pick the candidate signal and a held-out task' through 'A weight change is a proposed diff for a human to review', with explicit validation checkpoints throughout: CV as 'the noise-magnitude diagnostic, not a pass/fail gate', the combined-sum correlation bar ('Judge a candidate on the combined-sum delta, not the standalone figure'), the replication gate ('Only propose... once the SAME candidate clears the correlation bar on a second, different held-out task'), and error-recovery paths ('if no fresh task exists yet, say so explicitly and treat the result as weaker evidence'; 'If the runtime has no sub-agent dispatch available at all, stop and say so'). This matches the anchor-5 shape: clear sequence, explicit validation, feedback loops for failure modes.

5 / 5

Progressive Disclosure

The body is organized into clear sections (incident, 'What to do' steps, 'What this does NOT do', 'Related') and the Related section lists well-signaled one-level-deep references with markdown links, e.g. '[`../../eval/token_ceiling_formula.py`]... — the module under test', '[`../../eval/tuning/overfitting_guard.py`]... — HOLDOUT_TASKS'. No nested-reference chains exist, and no bundle files are provided to score against. It is not a 5 because the ~30-line incident narrative largely duplicates the linked results write-up ('2026-08-22-second-signal-experiment-genuinely-blind.md') and could live in a reference file with a two-line summary in the body.

4 / 5

Total

16

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states exactly what the skill does (blind multi-draw rating + correlation testing of candidate signals), the key methodological constraints (independent dispatches, replication on a second held-out task), and four explicit user trigger phrases. The only weakness is length — some methodology backstory in the middle sentences dilutes the action list, though it never becomes vague.

DimensionReasoningScore

Specificity

Quotes: 'Test whether a candidate real-work signal... deserves a nonzero default weight', 'Dispatches 3+ independent draws per held-out task, computes per-signal CV and Pearson correlation (alone, and added to the signal sum...)', 'requiring replication on a SECOND held-out task before proposing a nonzero weight'. These are multiple concrete, comprehensive actions covering the full workflow (test, dispatch, compute CV/correlation alone and combined, replicate, propose weight). It is not a 4 because there are no coverage gaps — every step of the method is named concretely; the explanatory asides ('which would turn blind rating into transcribing the answer key') are constraints, not vagueness.

5 / 5

Completeness

Quotes: 'Test whether a candidate real-work signal... deserves a nonzero default weight, via the blind multi-draw rating + correlation methodology' (clear what) and 'Use when someone says "test this new signal"...' (explicit when with concrete trigger phrases). Both what and when are answered explicitly and concretely, matching the anchor-5 example pattern; a 4 would require the 'when' to be merely present but less specific, which is not the case here.

5 / 5

Trigger Term Quality

Quotes: 'Use when someone says "test this new signal", "does [signal] deserve a nonzero weight", "re-run the signal validation experiment", or "validate real-work signals against real data"'. Four natural user phrasings is good keyword coverage for this niche domain. It is not a 5 because common variations like 'should this signal keep a zero weight', 'check the correlation', or references to the sibling tuning skills are absent, leaving a few natural trigger phrasings uncovered.

4 / 5

Distinctiveness Conflict Risk

Quotes: 'Test whether a candidate real-work signal in `eval/token_ceiling_formula.py`... deserves a nonzero default weight' and trigger phrases like 'validate real-work signals against real data'. This is a clear niche (signal weight validation for one specific formula module) with triggers that would not fire for the sibling skills (holdout tuning, prompt tuning) or any generic evaluation request. Minimal conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 12 suspicious

Warning

Total

14

/

16

Passed

Repository
Cloudzero/cloudzero-claude-marketplace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.