CtrlK
BlogDocsLog inGet started
Tessl Logo

launchdarkly-experiment-setup

Set up and run experiments in LaunchDarkly. Create experiments with metrics, treatments, and flag config, start iterations to collect data, swap design between iterations, and stop with a winner.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable skill with a clear lifecycle workflow, explicit verification steps, and strong error-recovery guidance. Weaknesses are minor: some padding and general experimentation principles, and inlined tool-reference material that could live in a separate reference file.

Suggestions

Trim the opening paragraph and "Core Principles" items that restate general A/B-testing knowledge Claude already has (e.g. "One change at a time").

Document `get-flag` / `get-flag-status-across-envs` in the tool list since Steps 2 and the edge-case table depend on them.

Consider moving the full MCP tool semantics and detailed JSON schemas into a single one-level-deep reference file to keep SKILL.md a lean overview.

DimensionReasoningScore

Conciseness

The body is largely efficient with dense JSON examples and tables, but the opening paragraph ("You're using a skill that guides you through...") and general principles like "One change at a time: test one variable per experiment" restate knowledge Claude already has — minor trimming opportunities, matching anchor 4 rather than the every-token-earns-its-place level.

4 / 5

Actionability

Copy-paste-ready JSON payloads for create, start, evolve, and stop plus a concrete edge-case table cover the common cases, but minor gaps remain: `get-flag` is referenced in Step 2 and the edge-case table without being documented in the tool list, and treatment/variation IDs must be discovered from responses without a shown example. Fits anchor 4, just short of fully-executable 5.

4 / 5

Workflow Clarity

A clear 7-step sequence with explicit validation (Step 5 confirms `currentIteration.status === "running"` and checks treatments/metrics), a genuine feedback loop for rejected `update-experiment` inputs (inspect `currentStatus`/`allowedFields`, then stop or use `save-and-start-experiment-iteration`), plus edge cases and a "What NOT to Do" checklist — matching the anchor for explicit validation with error-recovery loops. Not 4: checkpoints are present at every risky transition (create, start, mid-experiment changes, stop).

5 / 5

Progressive Disclosure

The single-file skill is well organized with clear sections (Prerequisites, Core Concepts, Workflow, Edge Cases) and no nested or dead references, but it runs ~230 lines with reference-style material (full MCP tool semantics, multiple JSON schemas) inlined that could be split into a reference file. Matches anchor 4; not 5 since a leaner SKILL.md pointing to one-level-deep references would be easier to navigate.

4 / 5

Total

17

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that names concrete lifecycle actions and is clearly distinguishable by the LaunchDarkly niche. Its main gaps are a missing explicit "Use when..." trigger clause and the absence of common synonyms like "A/B test".

Suggestions

Add an explicit trigger clause, e.g. "Use when the user wants to set up, run, or analyze A/B tests or experiments in LaunchDarkly, or asks about experiment metrics, treatments, or declaring winners."

Include natural synonyms users would say, such as "A/B test", "feature flag experiment", or "experimentation", to improve trigger-term coverage.

DimensionReasoningScore

Specificity

The description lists five concrete actions covering the full experiment lifecycle — "Create experiments with metrics, treatments, and flag config, start iterations to collect data, swap design between iterations, and stop with a winner" — which matches the comprehensive-coverage anchor; it is not the minor-gaps level of 4.

5 / 5

Completeness

The "what" is fully explicit (full lifecycle from creation through declaring a winner), but there is no "Use when..." clause or equivalent explicit trigger guidance anywhere in the description — per the rubric guideline this caps completeness at 3, exactly matching the anchor "clear 'what' but 'when' is missing or only weakly implied". Not 4: no when-guidance exists at all, so it cannot be 'more explicit yet present'.

3 / 5

Trigger Term Quality

Good keyword coverage ("experiments", "LaunchDarkly", "metrics", "treatments", "flag config", "winner"), but common natural variations such as "A/B test" or "feature flag testing" are missing, so it falls just short of the comprehensive-synonyms anchor.

4 / 5

Distinctiveness Conflict Risk

"LaunchDarkly" plus experiment-specific triggers (metrics, treatments, iterations, winner) define a clear niche with minimal overlap risk against generic analytics or feature-flag skills.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
launchdarkly/ai-tooling
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.