CtrlK
BlogDocsLog inGet started
Tessl Logo

launchdarkly-metric-choose

Choose the right metrics for a LaunchDarkly experiment, guarded rollout, or release policy. Use when the user wants to know which metrics to use, which is the primary metric for an experiment, what guardrails to add, or which events to monitor in a rollout. Surfaces what will auto-attach from existing release policies before making additional recommendations.

85

1.16x
Quality

81%

Does it follow best practices?

Impact

92%

1.16x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured advisory workflow that surfaces genuinely non-obvious LaunchDarkly domain knowledge and gives concrete tool calls, tables, and an output template. The main gap is the absence of explicit validate→fix→retry feedback loops and some prose that could be tightened.

Suggestions

Add an explicit verification checkpoint after Step 3 (e.g., 'If the user's intended primary metric is at-risk, stop and require instrumentation before proceeding') to convert the health classification into a hard gate.

Tighten explanatory prose such as 'Guarded rollouts are safety mechanisms, not experiments' into directive guidance to improve conciseness.

Parameterize the example tool calls (e.g., show the full argument shape for list-release-policies) so the guidance reads as executable rather than illustrative.

DimensionReasoningScore

Conciseness

Efficient overall — it assumes Claude's competence and avoids explaining what LaunchDarkly or experiments are, instead surfacing non-obvious domain knowledge (CUPED/percentile incompatibility, context-kind mismatches, auto-attach behavior). Not a 5 because several prose passages (e.g. 'Guarded rollouts are safety mechanisms, not experiments') restate reasoning that could be tightened.

4 / 5

Actionability

Provides concrete MCP tool calls (`list-release-policies(projectKey)`, `list-metrics`, `list-metric-events`), typed decision tables, and a copy-shaped output template. Not a 5 because the tool calls are illustrative rather than fully parameterized/executable and the skill is advisory by design, so guidance stops short of copy-paste-ready commands.

4 / 5

Workflow Clarity

Steps 1–5 are clearly sequenced with a context-split (experiment vs guarded rollout vs release policy) and a health-classification checkpoint in Step 3 before recommending. Not a 5 because validation is more classification than an explicit validate→fix→retry feedback loop, and the advisory (non-destructive) nature means no hard error-recovery checkpoints.

4 / 5

Progressive Disclosure

Well-organized into clear sections (Prerequisites, Workflow, Important Context, Related Skills) with one-level-deep pointers to sibling skills at the end. No bundle files exist (references/scripts/assets absent) and the body is ~175 lines, so it is not the under-50-line simple-skill exception; it is appropriately self-contained but slightly longer than a pure overview.

4 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description that explicitly covers what and when with concrete trigger phrases and a distinctive LaunchDarkly-specific niche. Minor room to broaden action variety and trigger-term synonyms.

DimensionReasoningScore

Specificity

Names the domain (LaunchDarkly experiment, guarded rollout, release policy) and several concrete actions — 'Choose the right metrics', 'which is the primary metric', 'what guardrails to add', 'which events to monitor', 'Surfaces what will auto-attach'. Not a 5 because the action set is narrow (all variants of selecting/surfacing metrics) rather than a comprehensive list of distinct operations.

4 / 5

Completeness

Explicitly answers both what ('Choose the right metrics for a LaunchDarkly experiment, guarded rollout, or release policy') and when (a concrete 'Use when the user wants to know...' clause with multiple trigger phrases). Matches the 5 anchor: clearly and explicitly answers both with concrete trigger phrases.

5 / 5

Trigger Term Quality

The 'Use when the user wants to know which metrics to use, which is the primary metric... what guardrails to add, or which events to monitor' clause captures natural phrasings users would say. Not a 5 because it lacks synonyms/extensions and repeats 'metric' heavily without varied terminology a user might actually say.

4 / 5

Distinctiveness Conflict Risk

'LaunchDarkly experiment, guarded rollout, or release policy' carves a clear niche with distinct, domain-specific triggers unlikely to fire for unrelated skills. Minimal overlap risk; matches the 5 anchor for a clear niche with distinct triggers.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 suspicious

Warning

Total

15

/

16

Passed

Repository
launchdarkly/ai-tooling
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.