CtrlK
BlogDocsLog inGet started
Tessl Logo

ultraqa

Adversarial dynamic e2e QA workflow - generate hostile scenarios, test, verify, fix, report, and clean up

59

Quality

74%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/ultraqa/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a dense, highly actionable operating card with an exemplary validated cycle and unambiguous exit contract. Its weaknesses are a buried, unlinked external reference and inline reference-grade material (the CLI command block), plus a few cryptic compressed lines and one duplicated phrase.

Suggestions

Move the full `omx state write/read/clear` command block (and optionally the 8-class scenario matrix) into a reference file and link it clearly, keeping only the two or three most-used commands in SKILL.md.

Fix the duplicated '**Fixes applied** / **Fixes applied**' line in the evidence/output contract.

Rewrite compressed lines like 'Keep outcome-first framing, local overrides for the active workflow branch, and `continue` on the current verified next step' into plain instructions, and signal `templates/AGENTS.md` with a proper link and a one-line description of what it contains.

DimensionReasoningScore

Conciseness

The body is lean and telegraphic with no explanations of concepts Claude already knows — every section carries operational content. It is not a 5 because a few compressed lines ('Keep outcome-first framing, local overrides for the active workflow branch...') are cryptic, and there is a duplicated phrase ('**Fixes applied** / **Fixes applied**') that is pure noise.

4 / 5

Actionability

It provides copy-ready commands (the full `omx state write/read/clear` block with exact JSON payloads), concrete flag-to-goal mapping, and executable patterns like `env -u OMX_ROOT -u OMX_STATE_ROOT` and `pathToFileURL(join(repoRoot, "dist", ...)).href`. Minor gaps ('safe substitutes', 'bounded CLI/service harness' left undefined) keep it below the fully-executable anchor.

4 / 5

Workflow Clarity

The 5-step cycle is clearly sequenced with an explicit validation gate ('CHECK RESULT: pass only when baseline, adversarial scenarios, evidence, and cleanup all pass. Otherwise diagnose and fix, then repeat'), explicit stop conditions (three repeats, max cycles), and a complete exit-status contract. This matches the anchor with feedback loops and checklists for a complex process.

5 / 5

Progressive Disclosure

Sections are well-organized, but the only external reference (`templates/AGENTS.md`) is buried in prose with no link or navigation, and the full CLI state-command block and the 8-class scenario matrix are inline content that plausibly belongs in a separate reference file. This matches 'some structure but could be better organized; references present but not clearly signaled'.

3 / 5

Total

16

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is compact and action-dense with a distinctive niche, but it omits any explicit 'when to use' trigger guidance and relies on abbreviated jargon ('e2e') rather than the natural phrases a user would say. Adding a 'Use when...' clause with spelled-out synonyms would materially raise completeness and trigger-term quality.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user asks for adversarial, hostile, or end-to-end QA of a runnable behavior, or when shallow test/lint/build checklists are not enough.'

Spell out 'end-to-end' alongside 'e2e' and include natural synonyms like 'testing', 'integration test', or 'QA' to improve trigger-term coverage.

Differentiate a couple of the generic verbs (test, verify, fix) with what they concretely cover, e.g. 'runs baseline tests plus hostile input, prompt-injection, and flaky-test scenarios' to push specificity toward 5.

DimensionReasoningScore

Specificity

The description lists six concrete actions ('generate hostile scenarios, test, verify, fix, report, and clean up') covering the full QA lifecycle, matching the 'several specific actions with minor gaps' anchor. It falls short of 5 because the verbs are terse and generic compared to comprehensively differentiated actions.

4 / 5

Completeness

It has a clear 'what' (adversarial dynamic e2e QA with enumerated actions) but no 'Use when...' clause or equivalent explicit trigger guidance, which per the judging guidelines caps completeness at 3.

3 / 5

Trigger Term Quality

Relevant keywords exist ('adversarial', 'e2e', 'QA', 'hostile scenarios'), but common natural variations users would actually say are missing: 'end-to-end', 'testing', 'integration test'. This matches 'some relevant keywords but missing common variations or synonyms' rather than the good-coverage anchor.

3 / 5

Distinctiveness Conflict Risk

'Adversarial dynamic e2e QA' with hostile scenario generation carves a distinct niche that is unlikely to trigger for ordinary build/lint/test requests, though it has minor overlap with general QA/testing skills — matching 'mostly distinct; minor overlap risk'.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Yeachan-Heo/oh-my-codex
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.