CtrlK
BlogDocsLog inGet started
Tessl Logo

creating-experiments

Guides agents through experiment creation: reading the project's setup with experiment-setup-context, defining the hypothesis, configuring rollout and bucketing, setting up analytics and running time, and reporting which choices are guesses. Delegates rollout decisions to configuring-experiment-rollout and metric setup to configuring-experiment-analytics. TRIGGER when: user asks to create a new experiment or A/B test, OR when you are about to call experiment-create. DO NOT TRIGGER when: user is updating an existing experiment, managing lifecycle, or only browsing experiments.

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

—

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-sequenced, actionable workflow with strong validation checkpoints and lean domain-specific guidance. Its one real defect is a broken bundle reference (references/setup-decisions.md is missing), which undermines progressive disclosure.

Suggestions

Create references/setup-decisions.md (the fact-to-choice/tier mapping the body relies on in Step 0) or remove the reference and inline the essential mapping so the agent is never sent to a missing file.

If keeping the reference, confirm the path is correct relative to the skill bundle and that the file ships alongside SKILL.md.

Consider surfacing the sibling-skill loads (configuring-experiment-rollout, configuring-experiment-analytics) as explicit 'load this skill before proceeding' links alongside the existing prose so delegation is unambiguous navigation.

DimensionReasoningScore

Conciseness

Largely lean and instruction-focused; the domain-specific detail (the two rollout_percentage meanings, deprecated 'parameters' keys, ensure_experience_continuity ordering) is non-generic knowledge Claude lacks rather than padding, with only minor trimmable explanation around the 'web' type.

4 / 5

Actionability

Provides a copy-paste-ready 'experiment-create' JSON payload covering the common case, plus concrete decision rules (bucketing fallback order, variant keying convention) and specific next actions for each step.

5 / 5

Workflow Clarity

Clear Step 0–3 sequence with explicit validation checkpoints ('Confirm both exist with read-data-schema', 'Never call a tool you can't see'), a branching fallback when setup-context is unavailable, and a three-group review report acting as a feedback/checklist step; launch is gated to avoid a destructive unprompted action.

5 / 5

Progressive Disclosure

Sectioning and sibling-skill delegation are well signaled, but the body instructs the agent to 'Apply references/setup-decisions.md to the result' and no references/ directory or setup-decisions.md file exists — a dead reference that breaks navigation, keeping it below the 'minor organization gaps' level.

3 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concrete, comprehensive, and uses natural trigger terms with explicit positive and negative trigger guidance, all in third person. It clearly distinguishes creation from sibling lifecycle/analytics/rollout skills.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'reading the project's setup', 'defining the hypothesis', 'configuring rollout and bucketing', 'setting up analytics and running time', 'reporting which choices are guesses' — with comprehensive coverage and named delegate skills.

5 / 5

Completeness

Explicitly answers both what (the full creation flow and its delegated sub-skills) and when (concrete 'TRIGGER when' / 'DO NOT TRIGGER when' phrases with examples), matching the top anchor.

5 / 5

Trigger Term Quality

Includes natural user phrasing — 'create a new experiment or A/B test' — alongside an explicit 'TRIGGER when' clause; not merely technical jargon, though 'experiment-create' is tool-specific it is appropriately scoped.

5 / 5

Distinctiveness Conflict Risk

Clear niche (new experiment creation) with explicit negative triggers excluding updates, lifecycle, and browsing, plus delegation to sibling skills, giving minimal conflict risk.

5 / 5

Total

20

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing

Warning

referenced_paths_exist

Referenced path issues: 2 missing

Warning

Total

14

/

16

Passed

Repository
PostHog/posthog
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.