Test whether a candidate real-work signal in `eval/token_ceiling_formula.py` (e.g. `context_ingestion_volume`, `investigative_uncertainty`, or new) deserves a nonzero default weight, via the blind multi-draw rating + correlation methodology this repo's research converged on. Ratings must come from genuinely independent sub-agent dispatches seeing ONLY a forward-looking task spec and the signal definitions — never a context holding the real actuals or this repo's retired write-ups, which would turn "blind rating" into transcribing the answer key. Dispatches 3+ independent draws per held-out task, computes per-signal CV and Pearson correlation (alone, and added to the signal sum — dilution, not weak correlation, is the dominant failure found twice), requiring replication on a SECOND held-out task before proposing a nonzero weight. Use when someone says "test this new signal", "does [signal] deserve a nonzero weight", "re-run the signal validation experiment", or "validate real-work signals against real data".
68
85%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
This skill hasn't been evaluated yet
f539a8b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.