CtrlK
BlogDocsLog inGet started
Tessl Logo

adversarial-ux-test

Roleplay a hostile user to find and triage UX pain points.

60

Quality

71%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/dogfood/adversarial-ux-test/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplar of an instruction-only skill: a crisply sequenced six-step workflow with a mandatory validation filter, concrete decision rules, templates, and worked persona examples, at the cost of only minor verbosity. Its main structural weakness is that everything lives in one ~180-line file when a persona library or tips could be offloaded to references.

DimensionReasoningScore

Conciseness

The body is almost entirely non-obvious instruction (persona questions, filter criteria, rant template, tips) with only minor trimmable padding such as 'Think of it as an automated "mom test" — but angry' and the motivational framing in 'Why This Works', matching the 'efficient; minor instances of over-explanation' anchor.

4 / 5

Actionability

Guidance is fully executable: a copy-paste rant template, explicit decision rules (filter criteria, 'more than 5 clicks... almost always a RED finding'), a concrete good/bad persona example, an industry persona table, and concrete ticket specs — specific examples cover the common cases.

5 / 5

Workflow Clarity

Six clearly sequenced steps with an explicit, mandatory validation checkpoint (the Step 4 pragmatism filter before any ticket creation), feedback loops ('If the persona has zero complaints... make them older'), and a closing Rules checklist — matching the top anchor for sequence, checkpoints, and checklists.

5 / 5

Progressive Disclosure

A single well-sectioned SKILL.md with no bundle files and clear headers, tables, and navigation; however, some inline content (the example-personas-by-industry table and the long Tips list) could be split into a reference file, so structure is good but not optimally split — matching anchor 4 rather than 5.

4 / 5

Total

18

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concise, distinctive description that clearly states what the skill does, but it lacks any 'when to use' trigger guidance and misses natural trigger phrases like 'usability' or 'user testing'. Adding a 'Use when...' clause would materially raise completeness and trigger term quality.

Suggestions

Append an explicit trigger clause, e.g. 'Use when the user asks for a UX review, usability test, user testing, or wants to find friction/pain points in their app.'

Broaden trigger vocabulary with natural synonyms users would say: 'usability', 'user testing', 'grumpy user test', 'UX audit'.

Mention one or two more concrete outcomes (e.g. 'browse the app as that persona and create prioritized tickets') to improve specificity.

DimensionReasoningScore

Specificity

The description names the domain ('UX pain points') and 2-3 concrete actions ('Roleplay a hostile user to find and triage'), but coverage is not comprehensive — it omits key capabilities like testing a live app, screenshots, and ticket creation, so it sits at the '1-2 concrete actions, but not comprehensive' anchor rather than 4.

3 / 5

Completeness

The 'what' is clear (roleplay a hostile user to find and triage UX pain points), but there is no 'Use when...' clause or equivalent explicit trigger guidance, capping completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

'hostile user' and 'UX pain points' are relevant keywords, but common natural variations a user would actually say — 'usability test', 'user testing', 'UX review' — are missing, matching the 'some relevant keywords but missing common variations' anchor.

3 / 5

Distinctiveness Conflict Risk

The adversarial-hostile-user framing carves a fairly distinct niche with minimal conflict risk against most skills; only minor overlap with closely related skills like general UX review or dogfooding, matching the 'mostly distinct; minor overlap risk' anchor.

4 / 5

Total

13

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.