CtrlK
BlogDocsLog inGet started
Tessl Logo

experiment-results-planning

Use when designing experiments, result tables, mock planning data, evaluation protocols, or results sections before real data are final

71

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

87%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A concise, highly actionable planning skill with concrete file contracts, templates, and a clear stage-gate sequence. The main gap is the absence of explicit error-recovery loops describing what happens when a gate fails.

Suggestions

For each stage gate, add a short "if this gate fails" recovery action (e.g., "if D1 traceability is incomplete, revise the protocol or mark the unsupported contribution as a limitation and re-run D1").

Add an explicit validate-fix-retry note for the figure handoff and decontamination (D4) steps, since those produce artifacts that can be re-checked.

DimensionReasoningScore

Conciseness

Lean bullet/table structure with no explanation of concepts Claude already knows; every section adds skill-specific guidance (file contracts, naming rules, gate artifacts) rather than padding.

3 / 3

Actionability

Gives concrete, copy-paste-ready guidance: exact file paths (plan/experiment-protocol.md, tables/table-schema.md), table templates with column headers, naming rules (mock_ prefix, [待真实实验替换] marker), and fill-in prose patterns.

3 / 3

Workflow Clarity

Stage gates D0-D5 provide an explicit sequenced checklist with required artifacts, but the workflow lacks error-recovery feedback loops — it never states what to do when a gate fails — which anchor 3 expects.

2 / 3

Progressive Disclosure

No bundle files exist and none are needed; the body is a self-contained, well-sectioned doc (Hard Gate, Protocol, Gates, Traceability, Mock Boundary, Table Schema, Figure Handoff, Prose Pattern) with clear navigation and no nested references.

3 / 3

Total

11

/

12

Passed

Description

85%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, concrete, trigger-first description with an explicit Use-when clause and a distinct niche. Its only weakness is trigger-term naturalness: some phrasing leans toward internal jargon over user-voiced terms.

Suggestions

Add user-natural trigger variants (e.g., "writing the results section", "planning ablations", "designing baselines and metrics") alongside the more technical terms.

Consider replacing or glossing "mock planning data" with phrasing a user would actually say, such as "placeholder or mock result tables for planning".

DimensionReasoningScore

Specificity

Lists multiple concrete actions/objects — "designing experiments, result tables, mock planning data, evaluation protocols, or results sections" — matching the anchor for several specific concrete actions rather than vague language.

3 / 3

Completeness

Explicitly states the action ("designing experiments, result tables, ...") and the trigger ("Use when ... before real data are final"), answering both what and when with an explicit Use-when clause.

3 / 3

Trigger Term Quality

Natural terms like "experiments", "result tables", and "results sections" are present, but "mock planning data" and "evaluation protocols" are more internal jargon than phrasing a user would naturally voice, and common variations are missing.

2 / 3

Distinctiveness Conflict Risk

Occupies a clear niche — pre-data experiment/result planning — with distinct triggers unlikely to fire for unrelated skills.

3 / 3

Total

11

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Norman-bury/research-writing-skill
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.