CtrlK
BlogDocsLog inGet started
Tessl Logo

experimentation-platform-orchestrator

A platform decision framework for experimentation. When to use Statsig vs PostHog vs GrowthBook vs Optimizely vs Amplitude vs Eppo vs Kameleoon. How to migrate between them. How to coordinate when multi-platform is genuinely warranted. The decisions that compound for years and the ones you can defer. Triggers on which experimentation platform, choose Statsig vs PostHog, evaluate experimentation tools, switch experimentation platform, migrate from Optimizely, consolidate experimentation tools, multi-platform experimentation, experimentation platform decision, ab test platform selection, feature flag platform vs experiment platform, warehouse-native experiments, vendor lock-in experimentation. Also triggers when a team is asking about cost, governance, or migration cost across experimentation tools, or when an evaluation is starting.

72

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A substantive, well-organized decision playbook that pushes detail into seven real reference files and gives a clear evaluation sequence plus concrete per-platform profiles and migration patterns. Its main weakness is mild verbosity in the framing prose and decision-guidance that is heuristic rather than directly executable.

Suggestions

Tighten the opening framing paragraphs ('looks easy at the start and compounds for years afterward', 'This skill is the discipline that makes the decision well the first time') which restate the motivation rather than add new actionable information.

For migrations — destructive/batch operations — add an explicit validate->fix->retry checkpoint loop (e.g., 'validate experiment results match on both platforms; if they diverge beyond X%, fix metric definitions and re-run') to push workflow_clarity to 5.

Consider moving the duplicated consideration lists (7 vs 11 considerations) into a single canonical list or reference file to avoid restating overlapping frameworks twice in the body.

DimensionReasoningScore

Conciseness

The body is mostly efficient prose that adds genuine domain knowledge (vendor-specific gotchas, pricing shapes, migration effort estimates) Claude does not already know, but it is long and includes some padded framing ('Picking an experimentation platform is one of those decisions that looks easy at the start and compounds for years afterward') that could be trimmed. Above the score-3 midpoint because the bulk is substantive, not concept recap.

4 / 5

Actionability

Provides concrete decision guidance: a 7-consideration and 11-consideration framework, per-platform Strengths/Gotchas/Ideal-customer profiles, a decision-matrix summary with context-to-platform mappings, and migration patterns with engineer-week effort estimates. Not score-5 because it offers decision heuristics rather than copy-paste executable commands, and the per-platform verification probes ('ask what variance estimator do you use for ratio metrics') are guidance rather than script.

4 / 5

Workflow Clarity

Sequences the evaluation clearly (answer 7 questions honestly first, then read per-platform profiles, then consult the decision matrix) and the migration section gives ordered rules (migrate experiments before flags, retire old platform only after in-flight experiments complete) with parallel-run/cut-over checkpoints. Not score-5 because validation checkpoints are present for migrations but less explicit for the selection decision itself, and there is no validate-then-retry loop spelled out for the destructive consolidation case.

4 / 5

Progressive Disclosure

Well-structured overview with one-level-deep references that are all real files verified in ./references/ (platform-decision-matrix, migration-playbook, mcp-capability-comparison, multi-platform-orchestration, cost-and-pricing-models, governance-and-team-setup, common-mistakes), each clearly signaled inline and enumerated again in a dedicated 'Reference files' section with one-line descriptions — easy navigation matching the score-5 anchor.

5 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that names seven concrete platforms and three clear action areas (selection, migration, multi-platform coordination), backed by an exhaustive list of natural trigger phrases. It cleanly answers both what and when, and carves out a distinct niche adjacent to but non-overlapping with related skills.

DimensionReasoningScore

Specificity

Names the domain (experimentation platform decision) and lists multiple concrete actions: 'When to use Statsig vs PostHog vs...', 'How to migrate between them', 'How to coordinate when multi-platform is genuinely warranted', and 'The decisions that compound for years and the ones you can defer' — comprehensive coverage matching the score-5 anchor.

5 / 5

Completeness

Explicitly answers 'what' (a platform decision framework for experimentation: when to use each platform, how to migrate, how to coordinate multi-platform) AND 'when' via 'Triggers on [phrases]' plus 'Also triggers when a team is asking about cost, governance, or migration cost across experimentation tools, or when an evaluation is starting' — matches the score-5 anchor with concrete trigger phrases.

5 / 5

Trigger Term Quality

Extensive natural keyword coverage including synonyms and product names users actually say: 'choose Statsig vs PostHog', 'evaluate experimentation tools', 'switch experimentation platform', 'migrate from Optimizely', 'multi-platform experimentation', 'ab test platform selection', 'feature flag platform vs experiment platform', 'warehouse-native experiments', 'vendor lock-in experimentation' — matches the score-5 anchor with synonyms and specific platform names.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (experimentation platform decision/migration/coordination across seven named vendors) with distinct triggers unlikely to collide with experiment-design, experimentation-analytics, or feature-flagging skills, which the body explicitly boundaries against — minimal conflict risk matching the score-5 anchor.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
rampstackco/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.