CtrlK
BlogDocsLog inGet started
Tessl Logo

stat-result-validator

Validate statistical research outputs for formulation quality, method-to- problem alignment, theory presence, experimental evidence, fair comparison, artifact completeness, and final-claim consistency.

57

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./external/agents/stat_research_agent/skills/stat-result-validator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

72%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a lean, well-organized validation checklist that respects the token budget and structures itself cleanly, but its guidance is concrete yet not fully executable and its multi-step workflow lacks explicit intermediate validation checkpoints.

Suggestions

Add a short sequenced procedure with explicit validation checkpoints between stages (e.g., 'After theory checks pass, run experimental checks; on FAIL, return to the failing stage') to lift workflow_clarity.

Provide concrete, executable validation aids (a checklist script or specific commands/queries to run against the listed artifacts) so actionability reaches copy-paste readiness.

Tighten abstract checklist items like 'Baselines are meaningful' into concrete, checkable criteria (e.g., 'baselines share the same data split and evaluation protocol').

DimensionReasoningScore

Conciseness

The body is lean checklists and artifact paths with no padded explanations of concepts Claude already knows (e.g., it never explains what 'theory' or 'comparison' means), assuming Claude's competence throughout.

3 / 3

Actionability

It gives concrete checklists, specific artifact file paths, explicit blocking failures, and a PASS/WARN/FAIL verdict scheme, but offers no executable validation tooling and some items stay abstract ('Baselines are meaningful'), so it is concrete but incomplete rather than copy-paste ready.

2 / 3

Workflow Clarity

The chain 'formulation -> method -> theory -> experiment -> comparison' supplies an implicit sequence and a final verdict checkpoint, but there are no explicit per-step validation checkpoints or fix-and-retry feedback loops between stages.

2 / 3

Progressive Disclosure

The skill is a self-contained single file with no bundle files, well-organized into clearly headed sections and no nested references (the progress/<TOPIC_ID>/ paths are artifacts to validate, not reference files), so the well-organized-sections anchor applies.

3 / 3

Total

10

/

12

Passed

Description

60%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific about what it validates but omits any explicit trigger guidance for when Claude should use it, and its terms lean technical rather than natural. Adding a 'Use when...' clause with user-facing keywords would lift completeness and trigger_term_quality.

Suggestions

Append a 'Use when...' clause naming natural user phrases (e.g., 'Use when auditing statistical research results, checking whether claims are backed by theory and experiments, or reviewing research artifact completeness').

Soften technical jargon toward terms users actually say (e.g., 'research claims', 'experiment evidence', 'fair baselines') to improve trigger_term_quality.

Add a distinctiveness cue tying the skill to the full research chain (formulation -> method -> theory -> experiment -> comparison) so it does not collide with generic QA skills.

DimensionReasoningScore

Specificity

The description states one concrete action ('Validate statistical research outputs') across seven specific dimensions (formulation quality, method-to-problem alignment, theory presence, experimental evidence, fair comparison, artifact completeness, final-claim consistency), matching the anchor that lists multiple specific concrete actions.

3 / 3

Completeness

It clearly answers 'what' (validate research outputs for the listed dimensions) but provides no explicit 'when'/trigger guidance, so per the rubric guideline a missing 'Use when...' clause caps completeness at 2.

2 / 3

Trigger Term Quality

Terms like 'formulation quality', 'theory presence', and 'artifact completeness' are technical/abstract rather than natural user phrases; some relevant keywords (validation, comparison, claims) appear but common natural variations are missing.

2 / 3

Distinctiveness Conflict Risk

The statistical-research-validation niche is somewhat specific, but without explicit distinct triggers in the description it could still overlap with general QA/audit skills.

2 / 3

Total

9

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
aiming-lab/AutoResearchClaw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.