CtrlK
BlogDocsLog inGet started
Tessl Logo

message-test-designer

Use when the user asks to "test our messaging before we scale it", "design a message-market-fit panel", or "run a 5-second comprehension test on our new tagline"; produces a message-test design spec — hypothesis, panel and recruit criteria, comprehension / 5-second / message-market-fit (Wynter-style) protocols, stimulus set drawn from the canon, success thresholds, and a stop/revise decision rule — for the TALE Evaluate phase so the message is validated before any paid scale. It designs the test; it never runs the experiment or adjudicates a claim. Not for running the panel or A/B experiment — use send-experiment-designer or ad-test-designer; not for analyzing the results — use performance-analyzer; not for authoring the message itself — use message-system-architect. 消息测试/理解度测试/面板设计/五秒测试/消息市场契合

73

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable design skill with a clear sequenced workflow, validation checkpoints, feedback loops, and good progressive disclosure through a labeled reference list. Its one weakness is conciseness: the same scope-handoff exclusions are repeated across several sections.

Suggestions

Consolidate the scope-guard exclusions into one canonical statement (the Scope guard paragraph) and reference it from later sections instead of restating the full 'not for running the panel / A/B / authoring / adjudicating' list in Instructions, Reference Materials, and Next Best Skill.

Trim the repeated 'operation: propose request to registry-events.py' phrasing in steps 5, the Writes contract, and Save Results to a single defined reference once and reuse a short handle thereafter.

Reduce cross-link density in the prose body — several sibling-skill links appear in both inline scope statements and the Reference Materials list; keeping them in the reference list only would tighten the running text.

DimensionReasoningScore

Conciseness

Operational and largely non-generic, but the scope-guard exclusions ("does not run the panel / not for running the A/B experiment / not authoring the message") are restated across the scope guard, instructions steps 5–7, Reference Materials, and Next Best Skill, padding the body with repetition that could be tightened into a single canonical statement.

2 / 3

Actionability

For an instruction-only skill the guidance is concrete and actionable: a numbered 8-step procedure, a worked threshold example ("≥70% ... restate the core benefit unaided after 5 seconds"), exact memory paths, and a copy-paste-ready command (`python3 "${CLAUDE_PLUGIN_ROOT}/scripts/connectors/experiment.py" proportion --control <pref_A> <n> --variant <pref_B> <n>`).

3 / 3

Workflow Clarity

Steps 1–8 are clearly sequenced with explicit checkpoints: a NEEDS_INPUT stop-and-route in step 1, a claim-scan validation in step 5, the stop/revise feedback loop in step 6 (failed test routes back to message-system-architect, not more spend), and a "Done when" checklist.

3 / 3

Progressive Disclosure

The body is a concise overview with well-organized sections (Quick Start, Skill Contract, Data Sources, Instructions, Reference Materials, Next Best Skill) and a clearly signaled one-level-deep reference list (tale-benchmark.md, skill-contract.md, CONNECTORS.md, SECURITY.md, plus named sibling skills) rather than inlining that material; no bundle files exist locally to verify against.

3 / 3

Total

11

/

12

Passed

Description

100%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a strong, complete trigger specification: it states concrete capabilities, natural user-spoken triggers, both what and when, and explicit boundaries against adjacent skills in third person. No weaknesses worth lowering any dimension.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — "hypothesis, panel and recruit criteria, comprehension / 5-second / message-market-fit (Wynter-style) protocols, stimulus set drawn from the canon, success thresholds, and a stop/revise decision rule" — matching the score-3 anchor for specific concrete actions.

3 / 3

Completeness

Explicitly answers both what ("produces a message-test design spec — ...") and when ("Use when the user asks to ...") with explicit triggers, satisfying the score-3 anchor; voice is third person ("It designs the test").

3 / 3

Trigger Term Quality

Quotes natural phrases a user would actually say ("test our messaging before we scale it", "design a message-market-fit panel", "run a 5-second comprehension test on our new tagline") plus keyword tags, giving good coverage of natural trigger terms.

3 / 3

Distinctiveness Conflict Risk

Clear niche (pre-scale message validation) with distinct triggers and explicit "Not for ... use send-experiment-designer or ad-test-designer; ... use performance-analyzer; ... use message-system-architect" exclusions, making conflict with sibling skills unlikely.

3 / 3

Total

12

/

12

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 34 suspicious

Warning

Total

13

/

16

Passed

Repository
aaron-he-zhu/aaron-marketing-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.