CtrlK
BlogDocsLog inGet started
Tessl Logo

aris-result-to-claim

Use when experiments complete to judge what claims the results support, what they don't, and what evidence is still missing. Codex MCP evaluates results against intended claims and routes to next action (pivot, supplement, or confirm). Use after experiments finish — before writing the paper or running ablations.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced workflow skill with executable commands, explicit validation checkpoints, and honest routing rules. Its only weaknesses are minor: slight verbosity and two large blocks that could be extracted into reference files.

DimensionReasoningScore

Conciseness

The body is efficient and assumes Claude's competence — no explanations of known concepts — but the Step 5 wiki-update block and framing sentences ("Experiments produce numbers; this gate decides what those numbers mean") could be trimmed slightly.

4 / 5

Actionability

Guidance is fully executable: concrete W&B API and ssh commands, a complete copy-paste Codex prompt template, a structured output field list, and exact research_wiki.py commands with flags covering the common verdict cases.

5 / 5

Workflow Clarity

A clear five-step sequence with verdict-based routing (no/partial/yes), an explicit re-run feedback loop after supplementary experiments, a low-confidence handling rule, and a fallback when Codex MCP is unavailable — checkpoints are explicit, not implicit.

5 / 5

Progressive Disclosure

The single SKILL.md is well-sectioned with clear headers and appropriately inline workflow content, but at ~155 lines the large Codex-prompt and wiki-update blocks are candidates for one-level-deep reference files, keeping it just short of the ideal split.

4 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete third-person actions, explicit and repeated when-to-use guidance, and a distinct niche with low conflict risk. The only gap is modest synonym coverage in its trigger terms.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "judge what claims the results support, what they don't, and what evidence is still missing", "evaluates results against intended claims and routes to next action (pivot, supplement, or confirm)" — with comprehensive coverage of the skill's behavior in third-person voice.

5 / 5

Completeness

Both what and when are explicitly and concretely answered: the what is judging claim support and routing to pivot/supplement/confirm, and the when is stated twice with concrete trigger phrases ("Use when experiments complete", "Use after experiments finish — before writing the paper or running ablations").

5 / 5

Trigger Term Quality

Natural phrases like "experiments complete", "experiments finish", "before writing the paper", and "running ablations" are present, but a few common variations (e.g. "results are in", "analyze results") are missing, so it falls just short of the comprehensive-synonyms anchor.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche — post-experiment claim adjudication in a research pipeline — with triggers tied to experiment completion and paper writing that are unlikely to fire for unrelated skills.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
OpenLAIR/dr-claw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.