CtrlK
BlogDocsLog inGet started
Tessl Logo

ads-test

Design and evaluate paid-ad experiments with hypotheses, randomization units, sample-size and duration assumptions, guardrails, platform experiment tools, analysis, and decision rules. Use for A/B test, split test, experiment design, hypothesis, statistical significance, sample size, test duration, or experiment readout.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A token-efficient, well-sequenced experimental-design checklist that covers setup, readout, and anti-patterns without padding. Its main weakness is actionability: the sample-size/duration calculation and the required versioned JSON output are mandated but never specified, leaving Claude to fill in method details.

Suggestions

Add the sample-size/duration calculation method (e.g., baseline rate, MDE, power/alpha defaults, or a formula/reference) so step 3 is executable rather than just mandated.

Include a minimal example of the versioned JSON setup/readout schema referenced in step 7 so the required output shape is concrete.

Name the specific platform experiment tools to check in step 2 (e.g., Google Ads campaign drafts/experiments, Meta A/B test tool) instead of the generic "platform experiment tools".

DimensionReasoningScore

Conciseness

The 20-line body is lean and assumes Claude's competence — it names required design elements and pitfalls ("Do not repeatedly peek and stop on a favorable result, call underpowered noise a winner") without explaining any concepts Claude already knows, so every token earns its place.

5 / 5

Actionability

The guidance names concrete artifacts ("treatment, control, randomization unit, population, primary metric, guardrails, minimum effect, and stopping rule") but leaves key execution details missing: no method or formula for "Calculate sample and duration", no example of the required "versioned JSON" output format, and no named platform experiment tools. This fits the anchor for some concrete guidance but incomplete with missing key details, rather than mostly-executable guidance.

3 / 5

Workflow Clarity

A clear 7-step sequence with an explicit validation gate — "verify assignment integrity and data completeness before estimating effect and uncertainty" — plus pre-registration of analysis and thresholds. It stops short of the top anchor because there is no feedback loop describing what to do when the verification checks fail.

4 / 5

Progressive Disclosure

The skill is under 50 lines with no bundle files and no inlined content that belongs in separate references; a short, well-organized single-purpose body fully satisfies the simple-skill exception for progressive disclosure.

5 / 5

Total

17

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly covers both capability and trigger conditions with concrete, domain-specific language and no padding. Its only weaknesses are a few missing natural trigger synonyms and slight overlap with generic statistics terminology.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "hypotheses, randomization units, sample-size and duration assumptions, guardrails, platform experiment tools, analysis, and decision rules" — giving comprehensive coverage of the paid-ad experimentation domain in third-person voice, matching the anchor for multiple specific concrete actions.

5 / 5

Completeness

It clearly answers both questions: the "what" ("Design and evaluate paid-ad experiments with..." decision rules) and an explicit "Use for..." clause with concrete trigger phrases, matching the top anchor exactly.

5 / 5

Trigger Term Quality

Triggers include natural phrases users would say ("A/B test, split test, experiment design, statistical significance, sample size, test duration, experiment readout"), but common variations like "lift test", "holdout", "ad testing", or platform-specific terms are missing, fitting the anchor for good coverage with a few natural terms missing rather than comprehensive coverage.

4 / 5

Distinctiveness Conflict Risk

Paid-ad experimentation is a clear niche with distinct triggers, but generic statistics terms ("hypothesis", "statistical significance", "sample size") create minor overlap risk with general statistics or analytics skills, placing it at the "mostly distinct; minor overlap risk" anchor rather than minimal conflict risk.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
AgriciDaniel/claude-ads
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.