CtrlK
BlogDocsLog inGet started
Tessl Logo

stat-result-validator

Validate statistical research outputs for formulation quality, method-to- problem alignment, theory presence, experimental evidence, fair comparison, artifact completeness, and final-claim consistency.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./external/agents/stat_research_agent/skills/stat-result-validator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a concise, well-structured audit checklist with concrete artifact paths, per-stage checks, blocking-failure conditions, and a final verdict scheme. It is actionable and well-organized, lacking only an explicit fix-and-retry loop for failed verdicts.

Suggestions

Add an explicit feedback loop after the final-claim verdict (e.g. 'On FAIL: identify the broken link in formulation -> method -> theory -> experiment -> comparison, remediate, and re-run the affected checks').

Tighten subjective checks ('baselines are meaningful', 'comparable conditions') with one concrete acceptance criterion each so the audit is mechanical rather than judgment-based.

Consider a short 'How to run this audit' sequence at the top so the section order is unambiguously a procedure, not just a reference.

DimensionReasoningScore

Conciseness

The body is a lean set of checklists with no padding and no over-explanation of concepts Claude already knows; a few check items could be tightened further, keeping it just below the lean-and-efficient 5 anchor.

4 / 5

Actionability

It gives concrete, executable guidance — exact artifact file paths, explicit per-section check lists, and a PASS/WARN/FAIL verdict scheme — though a few checks ('baselines are meaningful', 'comparable conditions') are subjective rather than mechanical.

4 / 5

Workflow Clarity

Checks are sequenced by research stage with explicit blocking-failure callouts in the formulation and theory sections and a final-claim verdict gate, yielding clear checkpoints; it stops short of a 5 because there is no explicit fix-and-recheck feedback loop after a FAIL.

4 / 5

Progressive Disclosure

Content is well organized into clearly headed sections in a single self-contained file with no nested references and no bundle files needed; it is over 50 lines so does not qualify for the simple-skill 5 exception, but structure is solid.

4 / 5

Total

16

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and distinct, naming a comprehensive set of validation dimensions for statistical research outputs. Its main weakness is the absence of an explicit 'when to use' trigger clause, which caps completeness at 3.

Suggestions

Add a 'Use when…' clause naming concrete trigger situations (e.g. 'Use when reviewing a statistical research pipeline before finalizing claims, or when auditing formulation/theory/experiment consistency').

Vary the action verbs beyond 'Validate' (e.g. 'Audit', 'Check', 'Trace') to strengthen the specificity of capabilities.

Surface one or two natural user phrasings such as 'statistical sanity check' or 'quality gate' into the description text itself rather than only in metadata keywords.

DimensionReasoningScore

Specificity

Lists seven concrete validation targets (formulation quality, method-to-problem alignment, theory presence, experimental evidence, fair comparison, artifact completeness, final-claim consistency), but all hang off a single verb ('Validate'), so coverage is comprehensive yet action variety is limited rather than score-5.

4 / 5

Completeness

The 'what' is clear and detailed, but there is no 'Use when…' clause or equivalent explicit trigger guidance, which per the judging guidelines caps completeness at 3.

3 / 5

Trigger Term Quality

Natural domain terms are present ('Validate', 'statistical research outputs', 'formulation', 'theory', 'comparison', 'claims') that a researcher would plausibly say, but a few common phrasings (e.g. 'audit', 'sanity check') appear only in metadata keywords, not the description itself.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche (end-to-end statistical research-chain validation) unlikely to fire for unrelated skills; minor overlap risk with a generic 'audit' skill since no explicit when-clause narrows the trigger.

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
aiming-lab/AutoResearchClaw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.