CtrlK
BlogDocsLog inGet started
Tessl Logo

experimental-design

Design experiments and studies BEFORE data is collected — choosing a design, randomizing, blocking, and laying out treatment combinations so results are interpretable. Use whenever someone is planning a study, asks how to assign subjects/samples to groups, mentions randomization, blocking, stratification, controls, factorial or fractional-factorial designs, design of experiments (DOE), screening many factors, response-surface optimization, crossover or repeated-measures or split-plot designs, cluster/group randomization, Latin squares, plate layouts, batch/run-order effects, replication vs. pseudoreplication, or sequential/adaptive/group-sequential designs. Trigger even for informal phrasings like "how should I set up this experiment", "how do I avoid confounding", "what's the best way to test these 6 factors", or "assign these mice to conditions". For computing the sample size or power once the design is chosen, use statistical-power; for analyzing data already collected, use statistical-analysis.

75

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable skill body: executable seeded code, a decision tree, a mistakes checklist, and a sequenced workflow, with detail correctly offloaded to real one-level-deep reference and script files.

Suggestions

Add an explicit 'validate the schedule' checkpoint in the Workflow (e.g., call arm_balance(sched) and confirm balance before archiving) so batch generation has a visible validation step rather than one shown only in the code example.

Tighten the pseudoreplication worked example in 'The mistakes that ruin studies' to the load-bearing rule (replicate at the level the treatment is randomized) and defer the worked mice/cells illustration to references/design_types.md.

Consider moving the four 'Key references' citations into a reference file to keep SKILL.md as a pure overview, though this is a minor conciseness point.

DimensionReasoningScore

Conciseness

Dense, information-rich content that assumes Claude's competence (Fisher's principles stated tersely) with a few spots of mild over-explanation (e.g., the pseudoreplication worked example) that could be trimmed, fitting the 'efficient; minor instances' anchor rather than the fully lean anchor.

4 / 5

Actionability

Copy-paste-ready, seeded, executable code covering the common cases (block/stratified/cluster randomization; full/fractional/Plackett-Burman/CCD designs) returning real-unit layouts, plus an inline arm_balance sanity check, matching the fully-executable anchor.

5 / 5

Workflow Clarity

An explicit 8-step sequenced workflow with auditability (seed/archive) and a cross-reference for sample size; validation is present (arm_balance, seed archiving) but is implicit rather than called out as an explicit checkpoint in the workflow steps, so it sits at clear-sequence-with-minor-gaps rather than the explicit-validation anchor.

4 / 5

Progressive Disclosure

Clear overview with well-signaled one-level-deep references to references/ and scripts/ files that all exist on disk, a dedicated Resources section repeating them, and content appropriately split, matching the clear-overview-with-one-level-references anchor.

5 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A model description: it states concrete actions, enumerates both technical and informal trigger terms, explicitly separates 'what' and 'when', and carves out a distinct niche by routing sibling tasks to other skills.

DimensionReasoningScore

Specificity

Names multiple concrete actions ('choosing a design, randomizing, blocking, and laying out treatment combinations') with comprehensive coverage across the experimental-design domain, matching the anchor for listing several specific concrete actions.

5 / 5

Completeness

Explicitly answers 'what' ('Design experiments...BEFORE data is collected') and 'when' ('Use whenever someone is planning a study...') with concrete trigger phrases plus explicit boundary routing to statistical-power and statistical-analysis.

5 / 5

Trigger Term Quality

Comprehensive natural-term coverage including synonyms ('factorial or fractional-factorial', 'DOE', 'screening many factors') and informal user phrasings ('how should I set up this experiment', 'assign these mice to conditions'), matching the comprehensive-coverage anchor.

5 / 5

Distinctiveness Conflict Risk

Clear niche of pre-data-collection design with explicit delegation to adjacent skills (statistical-power, statistical-analysis), giving distinct triggers and minimal conflict risk.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
K-Dense-AI/scientific-agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.