CtrlK
BlogDocsLog inGet started
Tessl Logo

adversarial-ux-test

Roleplay the most difficult, tech-resistant user for your product. Browse the app as that persona, find every UX pain point, then filter complaints through a pragmatism layer to separate real problems from noise. Creates actionable tickets from genuine issues only.

63

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/dogfood/adversarial-ux-test/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable and well-sequenced, with concrete templates, an explicit pragmatism-filter checkpoint, and clean section organization. The main weakness is mild verbosity in the introductory framing and some repetition of the filter concept.

Suggestions

Trim the 'Why This Works' intro and 'Think of it as an automated mom test' framing to tighten conciseness toward the lean anchor.

Consolidate the pragmatism-filter explanation so it is introduced once rather than re-explained across Steps 3 and 4, reducing token overhead.

DimensionReasoningScore

Conciseness

Mostly efficient and action-oriented, but includes some unnecessary framing ('Think of it as an automated mom test — but angry', 'Most QA finds bugs. This finds friction') and re-explains the pragmatism-filter concept across Steps 3 and 4, so it could be tightened rather than earning the 'lean, every token earns its place' anchor.

2 / 3

Actionability

Provides highly concrete, copy-paste-ready guidance: five persona-generation questions, a friction-categories checklist, a full rant template (THE GOOD/BAD/UGLY), a RED/YELLOW/WHITE/GREEN classification with six filter criteria, and an explicit ticket format — matching the 'concrete, specific, copy-paste ready' anchor for an instruction-only skill.

3 / 3

Workflow Clarity

Clear six-step sequence (Persona → Browse → Rant → Pragmatism Filter → Tickets → Report) with an explicit validation checkpoint in Step 4 ('Step OUT of the persona. Evaluate each complaint') that gates ticket creation, plus constraints (max 10 tickets, screenshots required), matching the 'clear sequence with explicit validation steps' anchor.

3 / 3

Progressive Disclosure

No bundle files exist, so the self-contained SKILL.md is scored on organization; it is split into clear, well-signaled sections (Why, How to Use, Steps 1-6, Tips, Example Personas table, Rules), which per the scoring notes earns a 3 for a skill with no need for external references.

3 / 3

Total

11

/

12

Passed

Description

67%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and distinctive, naming concrete actions and carving out a clear niche. Its weakness is the absence of explicit 'Use when...' trigger guidance and limited natural trigger-term coverage, which cap completeness and trigger-term quality at 2.

Suggestions

Add an explicit 'Use when...' clause naming natural user phrases (e.g. 'Use when the user asks for a UX/usability test, wants to dogfood an app, or asks to find UX friction in a product').

Broaden trigger terms with common variations users actually say ('UX test', 'user testing', 'usability review') to lift trigger-term coverage from 2 to 3.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Roleplay the most difficult, tech-resistant user', 'Browse the app as that persona, find every UX pain point', 'filter complaints through a pragmatism layer', 'Creates actionable tickets' — matching the 'lists multiple specific concrete actions' anchor.

3 / 3

Completeness

Clearly answers 'what does this do' but provides no 'Use when...' clause or equivalent explicit trigger guidance, so per the judging guidelines completeness is capped at 2.

2 / 3

Trigger Term Quality

Contains relevant domain keywords ('UX pain point', 'tickets', 'persona') but lacks common natural variations a user would actually say (e.g. 'usability test', 'user testing', 'UX review'), so it lands at 'some relevant keywords but missing common variations' rather than full coverage.

2 / 3

Distinctiveness Conflict Risk

The adversarial-persona-plus-pragmatism-filter framing is a clear, distinctive niche unlikely to trigger for the wrong skill, matching the 'clear niche with distinct triggers' anchor rather than merely 'somewhat specific'.

3 / 3

Total

10

/

12

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.