CtrlK
BlogDocsLog inGet started
Tessl Logo

result-to-claim

Use when experiments complete to judge what claims the results support, what they don't, and what evidence is still missing. Codex MCP evaluates results against intended claims and routes to next action (pivot, supplement, or confirm). Use after experiments finish — before writing the paper or running ablations.

63

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/result-to-claim/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Operationally excellent: executable scripts, a rigorous validation-before-judgment pipeline, and fail-closed handling of reviewer unavailability. The weaknesses are structural and budgetary — ~100 lines of inline shell with a duplicated resolver chain, plus heavy reliance on external shared-references that are not shipped in this skill's bundle.

Suggestions

Extract the ARIS_REPO/evidence_check and research_wiki resolver chains into a single shared script (e.g., scripts/aris_resolve.sh) or one reference page, and call it from both Step 1.5 and Step 5 instead of duplicating the ~10-line resolution block.

Move the Step 1.5 pre-check bash block and the Step 5 wiki-update procedure into scripts/ or references/ files, keeping SKILL.md as a lean overview that states the policy (warn-and-skip) and points to the executable helper.

Either bundle the cited shared-references files with the skill or inline the minimal content each section depends on (e.g., the reviewer fallback chain from reviewer-routing.md), so the skill is self-contained when installed without the full ARIS repo.

DimensionReasoningScore

Conciseness

No padding that explains known concepts, but the body carries ~100 lines of inline shell (the Step 1.5 pre-check script and the Step 5 wiki procedure) and duplicates the same ~10-line ARIS_REPO resolver chain in both steps verbatim. Rationale-heavy asides ("the exact bug this closes", "a loop can drive, never acquit") add length without adding operational content. Not 4 because the duplicated resolver block and inlined scripts are more than minor trims.

3 / 5

Actionability

Mostly executable: a complete copy-paste bash block for the evidence pre-check, a full Codex MCP invocation with model/config/prompt template, exact research_wiki.py add_experiment/add_edge command lines with flags, and the verdict fields to parse. Not 5 because Step 3.5 and Step 5 mix prose-condition pseudocode ("if research-wiki/ exists:", "for each claim resolved by this verdict") with real code, leaving those sections as guided templates rather than runnable commands.

4 / 5

Workflow Clarity

Clear numbered sequence (1 → 1.5 → 2 → 3 → 3.5 → 4 → 5) with explicit validation checkpoints at every risky point: deterministic evidence pre-check before the Codex call, JSON-validity check on pre-check output, integrity-audit handling with confidence downgrade, EXP_NODE_OK gating of wiki edges, and a fail-closed reviewer fallback chain. Per-verdict routing (yes/partial/no) with a re-run loop for partial completes the feedback structure.

5 / 5

Progressive Disclosure

Section structure and signaled links are decent — eight shared-references are cited as one-level-deep markdown links — but no bundle directory (references/, scripts/, assets/) exists, and the ../shared-references/ files the body depends on are not present alongside the skill, so every reference dangles from this bundle's perspective. Additionally the long inline bash helpers would more naturally live in scripts/ or a reference file. Not 2 because the body itself is well-sectioned and references are clearly signaled, not buried.

3 / 5

Total

15

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete actions, an explicit two-fold trigger (when experiments complete / after experiments finish), and a distinct niche in the research pipeline. Third person throughout, no fluff. The only gap is modest keyword breadth (no data-source terms like W&B, no "verify/interpret results" synonyms).

DimensionReasoningScore

Specificity

Lists several concrete actions — "judge what claims the results support, what they don't, and what evidence is still missing", "evaluates results against intended claims and routes to next action (pivot, supplement, or confirm)". Not 5 because it omits parts of the skill's coverage (collecting results from W&B/logs, integrity audit, wiki updates); not 3 because far more than 1–2 actions are named.

4 / 5

Completeness

Explicitly answers both: what — judging claim support/evidence gaps and routing to pivot/supplement/confirm — and when, twice and concretely ("Use when experiments complete" and "Use after experiments finish — before writing the paper or running ablations"). Matches the anchor for a clear what AND when with concrete trigger phrases; not 4 since the when-clauses are already fully explicit.

5 / 5

Trigger Term Quality

Good natural keyword coverage: "experiments complete", "experiments finish", "claims", "results", "evidence", "before writing the paper", "running ablations" — phrases a researcher would actually say. Not 5 because common variations/synonyms like "W&B", "interpret results", or "verify claims" are missing; well above 3's 'some relevant keywords'.

4 / 5

Distinctiveness Conflict Risk

Clear niche — post-experiment claim adjudication with a named reviewer (Codex MCP) — that is unlikely to fire for unrelated skills. Minor overlap risk with sibling research-pipeline skills sharing the "results/claims/experiments" vocabulary (ablation-planner, proof-checker), which keeps it below 5.

4 / 5

Total

17

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 3 suspicious

Warning

Total

13

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.