CtrlK
BlogDocsLog inGet started
Tessl Logo

review-experiment

Use when reviewing experiment results against a declared hypothesis, primary metric, and decision rule.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, highly actionable read-only review skill with explicit labels, a copy-paste output template, and a quality-check checklist. The only gap is the absence of an explicit error-recovery feedback loop, which matters little given the read-only, no-external-action design.

Suggestions

Optionally add a brief "If inputs conflict or are missing" sub-step under the workflow showing the retry path (re-request the missing input rather than reconstructing it) to close the feedback-loop gap.

DimensionReasoningScore

Conciseness

The body is lean and efficient: purpose, required inputs, an 8-step workflow, an output template, guardrails, a quality check, and an example invocation, with no padding or explanation of concepts Claude already knows. Every token earns its place.

5 / 5

Actionability

Despite being instruction-only, the guidance is concrete and executable: numbered steps with explicit labels ("Missing", "Descriptive arithmetic", "INCONCLUSIVE", "Supplied confounders", "Unmeasured possibilities"), a copy-paste-ready output template, and a concrete example invocation grounding the common case.

5 / 5

Workflow Clarity

An 8-step sequence with an explicit Quality-check checklist and an INCONCLUSIVE decision gate gives clear checkpoints; not a 5 because there is no error-recovery feedback loop (validate -> fix -> retry), though that is less critical for a read-only review.

4 / 5

Progressive Disclosure

No bundle files exist and none are needed; the skill is a single well-organized SKILL.md with clearly labeled sections (Purpose, Required inputs, Workflow, Output format, Guardrails, Quality check, Example invocation) that is easy to navigate and appropriately self-contained.

5 / 5

Total

19

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, third-person description with an explicit "Use when..." trigger and a clearly distinguishable niche. It is strong overall but fuses its what and when into one clause and omits common synonyms like "A/B test".

Suggestions

Separate the capability from the trigger, e.g. "Reviews experiment results against a predeclared hypothesis, primary metric, and decision rule. Use when evaluating A/B test or experiment outcomes against a predeclared rule."

Add natural synonyms users say, such as "A/B test", "experiment outcome", "lift", or "significance check", to broaden trigger coverage.

DimensionReasoningScore

Specificity

"reviewing experiment results against a declared hypothesis, primary metric, and decision rule" names the domain plus three concrete review criteria, comparable to the score-4 anchor listing several specific actions; not a 5 because it is one action evaluated against criteria rather than multiple distinct actions.

4 / 5

Completeness

The explicit "Use when reviewing experiment results against a declared hypothesis, primary metric, and decision rule" supplies both a concrete trigger and an embedded what, but they are fused into a single clause rather than the cleanly separated capability-list-plus-triggers pattern of the score-5 anchor.

4 / 5

Trigger Term Quality

"reviewing experiment results" plus "hypothesis", "primary metric", and "decision rule" give good keyword coverage a user would naturally say; not a 5 because common synonyms like "A/B test", "experiment outcome", or "lift" are missing.

4 / 5

Distinctiveness Conflict Risk

The predeclared-hypothesis/primary-metric/decision-rule framing carves a clear niche with distinct triggers and minimal overlap risk with other skills.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
llaskin/AI-SDR-Skill-Pack
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.