CtrlK
BlogDocsLog inGet started
Tessl Logo

result-to-claim

Use when experiments complete to judge what claims the results support, what they don't, and what evidence is still missing. A secondary Codex agent evaluates results against intended claims and routes to next action (pivot, supplement, or confirm). Use after experiments finish — before writing the paper or running ablations.

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-engineered, highly actionable workflow with strong sequencing, validation checkpoints, and fail-closed error handling — its strongest qualities. It is weakened by a monolithic single-file structure with no bundle files, unverifiable cross-bundle references, pseudocode blocks that are not executable, and some repetition of the same caveats.

Suggestions

Extract the Step 5 wiki-update protocol and the Step 1.5 evidence pre-check into bundled reference or script files (e.g., references/wiki-update.md, scripts/evidence_precheck.sh) so the pseudocode becomes executable and SKILL.md stays an overview.

State the same-family/provisional acceptance caveat once (e.g., in a short 'Assurance model' section) and reference it from Step 1.5 and Review Tracing instead of repeating it three times.

Rewrite Step 3.5 and the routing-rule fallback chain as concrete runnable checks or a decision table rather than prose pseudocode, so every step is copy-paste executable like Steps 1–2.

DimensionReasoningScore

Conciseness

The body is dense and operational (W&B API call, fallback-resolving bash, a complete spawn prompt) with almost no explanation of concepts Claude already knows, fitting the level-4 'efficient; minor instances of over-explanation' anchor. Not level 5 because the same-family/provisional caveat is repeated three times (header note, Step 1.5, Review Tracing) and 'A single positive result on one dataset does not support a general claim' appears in both the prompt template and Rules — trimming would tighten it.

4 / 5

Actionability

Mostly executable: Step 1.5's bash pre-check is copy-paste ready, the spawn_agent prompt is fully specified, and routing/wiki steps give concrete commands. Falls short of level 5's 'fully executable, copy-paste ready' because Step 3.5 is pure pseudocode and Step 5's block mixes shell tests with non-shell control flow (`if [ "$EXP_NODE_OK" = 1 ]:`, `for each claim resolved by this verdict:`) that cannot run as written.

4 / 5

Workflow Clarity

Clear sequence (collect → pre-check → judge → parse → integrity check → route → wiki → trace) with explicit validation checkpoints (deterministic evidence pre-check, integrity audit pass/warn/fail handling, low-confidence treated as inconclusive) and feedback loops (re-run after supplementary experiments, multiple-partials escalation, fail-closed reviewer fallback chain with a traced BLOCKED record). Matches the level-5 anchor including error-recovery paths.

5 / 5

Progressive Disclosure

The skill ships no bundle at all (no references/, scripts/, or assets/), and its cited helper docs (`../shared-references/evidence-precheck.md`, `shared-references/experiment-integrity.md`, `../shared-references/review-tracing.md`, `reviewer-routing.md`) resolve outside the skill directory, so a reader cannot verify them from the bundle. Inlined ~50-line wiki-update pseudocode and the ~40-line fail-closed routing rule are content that belongs in separate reference/script files, matching the level-3 anchor 'content that should be separate is inline' better than level 4's 'most content is appropriately placed'.

3 / 5

Total

16

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete capabilities in third-person voice, gives two explicit 'when' trigger clauses with correct temporal positioning in the research workflow, and carves out a distinct niche. The only gap is modest synonym coverage for trigger phrases.

DimensionReasoningScore

Specificity

Quotes multiple concrete actions with comprehensive coverage: "judge what claims the results support, what they don't, and what evidence is still missing", "A secondary Codex agent evaluates results against intended claims and routes to next action (pivot, supplement, or confirm)". Matches the level-5 anchor ('lists multiple specific concrete actions; comprehensive coverage') and exceeds level 4 because the actions span the full pipeline (judge, evaluate, route with named outcomes).

5 / 5

Completeness

Explicitly answers both questions with concrete trigger phrases: what ("A secondary Codex agent evaluates results against intended claims and routes to next action") and when ("Use when experiments complete" and "Use after experiments finish — before writing the paper or running ablations"). Matches the level-5 anchor exactly; level 4 would require a weaker or less explicit 'when' clause.

5 / 5

Trigger Term Quality

Good natural-term coverage: "experiments complete", "results", "claims", "evidence", "paper", "ablations" — phrases a researcher would naturally say. Not level 5 because common synonyms and variations are missing (e.g., 'experiment results', 'interpret my results', 'validate findings', 'what can I claim'); above level 3 because the terms present are the natural phrasing, not jargon.

4 / 5

Distinctiveness Conflict Risk

Clear niche (post-experiment claim adjudication in a research pipeline) with distinct triggers ('after experiments finish', 'before writing the paper'). Minimal conflict risk; the mention of 'ablations' only names an adjacent routing target, not an overlapping capability, so it does not fall to level 4's 'minor overlap risk'.

5 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 suspicious

Warning

Total

13

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.