CtrlK
BlogDocsLog inGet started
Tessl Logo

octopus-ui-ux-design

Design UI/UX systems with style guides, palettes, typography, and component specs for new interfaces

57

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/octopus-ui-ux-design/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

66%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers an exceptionally actionable, well-sequenced pipeline with genuine validation feedback loops (preflight, dial range checks, contrast checker exit codes, adversarial critique). It is undermined by severe padding — enforcement-theater repetition and a 130-line inline persistence script — and by progressive-disclosure failures: referenced files that are not in the bundle and scripts inlined instead of split out.

Suggestions

Move the ~130-line persistence bash block into scripts/persist-design.sh and invoke it with arguments, cutting the body by roughly a quarter.

Collapse the repeated MANDATORY/HARD-GATE/You-MUST framing into a single execution-contract note at the top; the steps themselves already encode the sequence.

Ship the referenced files (skills/blocks/design-taste.md, codex-host-adapter.md) inside the bundle or inline their essential checklists, so no instruction depends on a file the skill does not carry.

DimensionReasoningScore

Conciseness

The 596-line body is noticeably verbose: repeated enforcement shouting ('EXECUTION CONTRACT (MANDATORY - CANNOT SKIP)', 'You MUST', 'MANDATORY', 'Do NOT skip it' appear dozens of times), a 130-line fully-inlined bash persistence script with line-by-line failure handling, and a full emoji banner template — padded sections that assume Claude will disobey rather than leverage its competence. It is above anchor 1 (the content is operational, not explaining concepts Claude knows), but well below anchor 3's 'mostly efficient'.

2 / 5

Actionability

Nearly everything is copy-paste executable: the BM25 search invocations ('python3 "$SEARCH_PY" ... --domain product'), the preflight check with READY/MISSING states, the WCAG contrast checker ('contrast-check.py \"<text-hex>:<bg-hex>\"'), the provider-availability bash loop, a complete AskUserQuestion JSON template, and a fully-formed atomic persistence script including lock acquisition and cleanup traps. Anchor 5's 'fully executable; copy-paste ready' fits; only intentional placeholders like '<user\'s product description>' remain, which is appropriate.

5 / 5

Workflow Clarity

Steps 1-8 are clearly sequenced with explicit validation checkpoints and feedback loops: preflight 'Only continue to Step 4 when preflight returns READY'; dial values validated before every search call ('stop and obtain a valid value rather than clamp it'); 'Exit 1 means at least one pair fails WCAG AA — fix the palette and re-run before delivering'; and the critique step institutionalizes a revise loop ('Fix must-fix items... Show the diff'). This matches anchor 5's 'explicit validation steps; feedback loops for error recovery'.

5 / 5

Progressive Disclosure

No references/, scripts/, or assets/ directories exist, yet the body repeatedly points to files outside the skill ('skills/blocks/design-taste.md', 'skills/blocks/codex-host-adapter.md', 'skill-design-lineage') that are not in the bundle, and inlines ~130 lines of persistence bash that clearly belongs in a scripts/ file. This matches anchor 2's 'content that clearly belongs in separate files is inlined' — there is real section structure, but the dangling references and monolithic bulk keep it below anchor 3.

2 / 5

Total

14

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and reasonably distinct, naming concrete deliverables (style guides, palettes, typography, component specs) in tight third-person voice. Its main defect is the complete absence of a 'when to use' clause, which caps completeness and weakens its trigger guidance.

Suggestions

Add an explicit trigger clause, e.g. 'Use when creating a new interface, building a design system, or when the user asks for palettes, style guides, or component specs.'

Broaden trigger terms to cover natural synonyms users say — 'design system', 'design tokens', 'wireframes', 'look and feel' — to raise trigger-term coverage.

Mention the validation capability (WCAG contrast checks, accessibility audit) in the description so the 'what' covers what the skill actually does end-to-end.

DimensionReasoningScore

Specificity

The description lists several concrete actions — 'style guides, palettes, typography, and component specs' — naming specific deliverables rather than generic 'helps with design' language. It stops short of anchor 5's comprehensiveness because it omits common deliverables it actually performs (design tokens, page layouts, contrast/accessibility validation), so 'minor gaps in coverage' (anchor 4) is the best fit; it is clearly above anchor 3, which expects only 1-2 concrete actions.

4 / 5

Completeness

The 'what' is clear (design UI/UX systems with the four listed deliverables) but there is no 'when' clause — no 'Use when...' or equivalent explicit trigger guidance anywhere in the description. Per the judging guidelines, a missing 'Use when...' clause caps completeness at 3; it is above anchor 2 because the 'what' is concrete, not vague.

3 / 5

Trigger Term Quality

Good natural keyword coverage: 'UI/UX', 'style guides', 'palettes', 'typography', 'component specs', 'interfaces' are all phrases a user requesting design work would plausibly say. A few natural terms are missing — no synonyms like 'design system', 'wireframes', 'look and feel', or 'design tokens' — matching anchor 4 rather than anchor 5's comprehensive synonym coverage, while clearly exceeding anchor 3's 'some relevant keywords'.

4 / 5

Distinctiveness Conflict Risk

'UI/UX systems', 'palettes', 'component specs' carve out a clear design niche that would not fire for document, data, or code skills. Minor overlap risk remains with closely related skills (e.g., general web-design or brand-typography skills), matching anchor 4's 'mostly distinct; minor overlap risk' rather than anchor 5's fully distinct trigger set.

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (603 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.