CtrlK
BlogDocsLog inGet started
Tessl Logo

visual-ralph

Visual Ralph orchestration for frontend UI from generated references, static references, or live URL targets, using $ultragoal with built-in visual verdict and pixel-diff evidence until the implementation matches and leaves a reproducible design system.

60

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./plugins/oh-my-codex/skills/visual-ralph/SKILL.md

The canonical home for this skill is visual-ralph in Yeachan-Heo/oh-my-codex

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, lean orchestration guide with a genuinely strong workflow: an explicit approval gate, a verdict-driven feedback loop with a numeric threshold, secondary diff evidence, and concrete completion/stop conditions. Its main weaknesses are unexplained placeholders and missing commands for screenshot capture and verdict invocation, plus redundancy between the workflow, completion checklist, and handoff template.

Suggestions

Provide (or reference) the concrete commands for capturing the current screenshot and invoking the Visual Ralph verdict, since step 5 depends on both but neither is executable as written.

Consolidate the overlap between the workflow steps, the 'Completion evidence and stop conditions' checklist, and the handoff template — e.g., keep the checklist as terse references to step numbers rather than restating each condition.

Move dense enumerations such as the URL-capture artifact contents in step 2 into a bundle reference file (e.g., references/url-capture.md) to keep SKILL.md a true overview.

DimensionReasoningScore

Conciseness

The body is lean and imperative with no padding and no explanation of concepts Claude already knows; every section instructs rather than describes. It falls short of anchor 5 because of deliberate redundancy — the 'Completion evidence and stop conditions' checklist restates workflow steps 3, 5, 6, and 7, and the handoff template repeats them a third time, which could be consolidated.

4 / 5

Actionability

Mostly executable guidance: a complete 'omx imagegen continuation' command with concrete paths, a copy-paste handoff template, an explicit threshold ('If score < 90'), and required verdict JSON keys ('score', 'verdict', 'category_match', 'differences[]', 'suggestions[]', 'reasoning'). Minor gaps keep it below 5 — no command is given for capturing screenshots or invoking the Visual Ralph verdict itself, and some placeholders (<session-id>, <slug-or-filename>) are unexplained.

4 / 5

Workflow Clarity

A clear 7-step sequence with an explicit approval-gate checkpoint ('Stop after generation or URL capture and obtain approval'), a validate-fix-retry feedback loop ('If score < 90, turn differences[] and suggestions[] into the next edit plan and rerun before editing'), secondary diff evidence, and a completion checklist with stop conditions and concrete-blocker reporting. This matches the top anchor: explicit validation steps, feedback loops, and a checklist.

5 / 5

Progressive Disclosure

Well-sectioned overview that correctly delegates shared invariants to templates/AGENTS.md with a clearly stated purpose ('Shared operating, delegation, state, hook, team, cancellation, and verification invariants live in...'). Below 5 because the referenced path lies outside this bundle (no references/, scripts/, or assets/ directories exist), leaving navigation unverifiable, and dense inline checklists (step 2's URL-capture artifact contents) could be split into a reference file.

4 / 5

Total

17

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates a concrete, distinctive pipeline (reference → $ultragoal implementation → visual verdict + pixel diff → design system) but is written from the system's perspective rather than the user's. It lacks any explicit 'Use when...' trigger clause and misses natural user synonyms like 'mockup', 'clone', or 'match a design', while over-indexing on internal jargon.

Suggestions

Append an explicit trigger clause, e.g., 'Use when the user wants a web/app UI built or restyled to match a mockup, screenshot, or live URL, or asks to clone a site's look.'

Include natural user synonyms such as 'mockup', 'screenshot', 'clone a website', and 'match a design' alongside the existing 'generated references' and 'live URL targets'.

Reduce internal jargon in the description ('Visual Ralph orchestration', '$ultragoal') in favor of capability-level phrasing the user would recognize.

DimensionReasoningScore

Specificity

The description lists several concrete capabilities — implementing 'frontend UI from generated references, static references, or live URL targets', 'built-in visual verdict and pixel-diff evidence', and leaving 'a reproducible design system' — but coverage has minor gaps (e.g., the approval gate and iteration loop are only implied). It is above anchor 3 because multiple specific actions are named, not just 1-2, but below anchor 5's comprehensive coverage.

4 / 5

Completeness

The 'what' is clearly stated (orchestrates reference/URL-driven UI implementation with visual verdict and pixel-diff evidence until match, producing a design system), but there is no 'Use when...' clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines. It is not lower because the 'what' is concrete and multi-part.

3 / 5

Trigger Term Quality

Some relevant natural keywords are present ('frontend UI', 'live URL', 'design system', 'references'), but common user phrasings like 'mockup', 'clone a website', 'match a design', or 'screenshot' are missing, and the description leans on skill-internal jargon ('Visual Ralph orchestration', '$ultragoal', 'pixel-diff'). This matches anchor 3 (relevant keywords but missing common variations/synonyms) rather than 4.

3 / 5

Distinctiveness Conflict Risk

The reference-matched implementation loop with verdict scoring and pixel-diff evidence is a distinct niche with minimal conflict risk against unrelated skills. It is not 5 because the absence of 'when' guidance leaves overlap risk with generic 'build or restyle a UI' requests, which broader frontend skills would also match.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 suspicious

Warning

Total

15

/

16

Passed

Repository
Yeachan-Heo/oh-my-codex
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.