CtrlK
BlogDocsLog inGet started
Tessl Logo

message-test-designer

Use when the user asks to "test our messaging before we scale it", "design a message-market-fit panel", or "run a 5-second comprehension test on our new tagline"; produces a message-test design spec — hypothesis, panel and recruit criteria, comprehension / 5-second / message-market-fit (Wynter-style) protocols, stimulus set drawn from the canon, success thresholds, and a stop/revise decision rule — for the TALE Evaluate phase so the message is validated before any paid scale. It designs the test; it never runs the experiment or adjudicates a claim. Not for running the panel or A/B experiment — use send-experiment-designer or ad-test-designer; not for analyzing the results — use performance-analyzer; not for authoring the message itself — use message-system-architect. 消息测试/理解度测试/面板设计/五秒测试/消息市场契合

71

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with a clear, validated workflow and clean progressive disclosure, but is let down by repeated scope-guard prose and jargon padding that could be tightened without losing meaning.

Suggestions

Consolidate the scope guard: state the 'designs only, hands off execution/analysis/authoring' boundary once instead of repeating it in the intro, the Scope Guard paragraph, and the Data Sources/Instructions sections.

Trim repeated protocol jargon (measurement-contract / evidence-observation / action-receipt) to a single defining mention plus a pointer to the binding reference.

Collapse the frontmatter description's exclusion list and the body's 'Not for…' lines into one canonical handoff list to remove duplicate tokens.

DimensionReasoningScore

Conciseness

Mostly efficient but padded: the scope guard and 'not for running/analyzing/authoring' exclusions are restated across the frontmatter, intro, and a dedicated Scope Guard paragraph, and protocol jargon is repeated.

3 / 5

Actionability

An 8-step procedure with concrete thresholds (≥70% restate), exact memory paths, and an executable experiment.py proportion command gives mostly executable guidance with only minor gaps.

4 / 5

Workflow Clarity

Clearly sequenced steps with explicit checkpoints: a NEEDS_INPUT stop, 'Done when' criteria, a failed-test stop/revise feedback loop, and termination/max-depth rules.

5 / 5

Progressive Disclosure

Clear overview with a single well-signaled one-level-deep reference (references/stimulus-binding.md, confirmed present and non-chaining) and clean section organization.

5 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description: concrete capabilities, natural trigger phrases, explicit when/what, and clear deconfliction from neighboring skills. Third-person voice is maintained throughout.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — hypothesis, panel/recruit criteria, comprehension/5-second/message-market-fit protocols, stimulus set from canon, success thresholds, and a stop/revise rule — giving comprehensive coverage.

5 / 5

Completeness

Explicitly answers both 'what' (produces a message-test design spec with listed components) and 'when' (Use when the user asks to…) with concrete trigger phrases.

5 / 5

Trigger Term Quality

Embeds natural user phrasings ('test our messaging before we scale it', 'design a message-market-fit panel', 'run a 5-second comprehension test on our new tagline') plus synonyms and Chinese keyword variants.

5 / 5

Distinctiveness Conflict Risk

States a clear niche (test design only) and explicitly routes adjacent work to send-experiment-designer, ad-test-designer, performance-analyzer, and message-system-architect, minimizing misfires.

5 / 5

Total

20

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 32 suspicious

Warning

Total

13

/

16

Passed

Repository
aaron-he-zhu/aaron-marketing-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.