CtrlK
BlogDocsLog inGet started
Tessl Logo

content-experimentation-best-practices

Content experimentation and A/B testing guidance covering experiment design, hypotheses, metrics, sample size, statistical foundations, CMS-managed variants, and common analysis pitfalls. Use this skill when planning experiments, setting up variants, choosing success metrics, interpreting statistical results, or building experimentation workflows in a CMS or frontend stack.

64

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/content-experimentation-best-practices/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A clean, well-organized overview that practices good progressive disclosure by deferring detail to four clearly-labeled reference files. Its weaknesses are explanatory padding of well-known concepts and the absence of any concrete sequenced workflow in the body itself.

Suggestions

Trim or remove the 'Core Concepts' section — definitions of A/B testing, statistical significance, and HiPPO are concepts Claude already knows and add token cost without value.

Add a brief, concrete decision workflow for selecting a reference (e.g. 'If designing a test → experiment-design.md; if unsure about significance → statistical-foundations.md') to replace the generic 'Start with the reference that matches'.

Surface one or two immediately actionable items inline (e.g. a sample-size heuristic or a minimal variant-config snippet) so the body is not purely an index.

DimensionReasoningScore

Conciseness

The body is short and mostly lean, but the 'Core Concepts' section explains basics Claude already knows ('Comparing two variants (A vs B) to determine which performs better', 'The confidence level that results aren't due to random chance') that could be trimmed.

3 / 5

Actionability

The body offers concrete navigation guidance (a matched reference list with one-line descriptions) but defers all executable how-to into references; as an instruction-only overview it is actionable for selection yet incomplete for execution.

3 / 5

Workflow Clarity

No sequenced multi-step workflow is present; the only process-like guidance is 'Start with the reference that matches the current problem', which is a useful selection heuristic but lacks explicit checkpoints or a clear sequence.

3 / 5

Progressive Disclosure

The body is a concise overview pointing to four real, one-level-deep reference files (verified present in references/), each clearly signaled with a descriptive label, making navigation easy.

5 / 5

Total

14

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description that clearly states both the skill's scope and its trigger conditions with concrete, natural-language phrases. Only minor: a few common synonyms (split testing, multivariate) are missing.

DimensionReasoningScore

Specificity

Lists multiple concrete capabilities and actions — 'experiment design, hypotheses, metrics, sample size, statistical foundations, CMS-managed variants' and 'planning experiments, setting up variants, choosing success metrics, interpreting statistical results' — giving comprehensive coverage rather than a single domain mention.

5 / 5

Completeness

Explicitly answers both what ('guidance covering experiment design, hypotheses, metrics...') and when ('Use this skill when planning experiments, setting up variants, choosing success metrics...') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Strong natural-term coverage with synonyms ('A/B testing', 'experiments', 'experimentation', 'success metrics', 'CMS'), but common variations like 'split testing' and 'multivariate testing' are absent from the description.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (content experimentation / A/B testing in a CMS or frontend stack) with distinct triggers and minimal overlap risk against general analytics or CMS skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
sanity-io/agent-toolkit
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.