CtrlK
BlogDocsLog inGet started
Tessl Logo

ulw-loop

A goal-like loop that decomposes work into systematic, evidence-bound ultrawork steps. Use when the user wants a goal loop or durable, checkpointed execution.

56

Quality

65%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./packages/omo-codex/plugin/components/ulw-loop/skills/ulw-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable overview with verified one-level-deep references and an explicit first-steps sequence. Its main weaknesses are time-sensitive model-tier/version details in a live section, some repetition in the Non-Negotiables, and detail (tool mapping, team mode) inlined that belongs in the reference bundle.

Suggestions

Move the model-tier mapping (GPT-5.6 sol/terra vs gpt-5.5/gpt-5.6-luna) into a versioned or 'compatibility' section in references/full-workflow.md so stale version numbers do not sit in the always-loaded body.

Push the Codex Tool Mapping table and the team-mode merge/conflict rules into a reference file, keeping only the decide-vs-default decision rule in the body.

Deduplicate the Non-Negotiables: goal registration is mandated in both Required First Steps and again in two separate bullets; consolidate into one authoritative statement.

DimensionReasoningScore

Conciseness

The body is dense and mostly information-rich, but it embeds time-sensitive model/version numbers ('GPT-5.6 (sol/terra)', 'gpt-5.6-luna', 'gpt-5.5') in a live section rather than a deprecated/old-patterns section, and several Non-Negotiables restate goal-registration rules. Not a 4 because the rubric explicitly penalizes version-number content outside an old-patterns section and the bullets could be tightened.

3 / 5

Actionability

Concrete executable guidance dominates: exact commands ('omo-agent-toolkit ulw-loop create-goals', 'status --json', '--session-id <id>'), error codes ('ULW_LOOP_SESSION_SCOPE_REQUIRED'), and full tool-call signatures in the mapping table. Not a 5 because several steps defer abstractly ('follow the full workflow's delegation and evidence rules') instead of giving the exact next command.

4 / 5

Workflow Clarity

'Required First Steps' is an explicit ordered sequence naming the reference sections to read, and evidence/gate rules ('Record only after cleanup receipts exist', 'gates are green', 'never relabel or regenerate') act as validation checkpoints. Not a 5 because the core execution loop and its feedback loop live in the reference file, leaving body-level checkpoints partial.

4 / 5

Progressive Disclosure

The body is an overview pointing to two real, one-level-deep references (full-workflow.md and define-goal.md, both verified to exist with the sections cited), each clearly signposted with what to read when. Not a 5 because the body inlines the full Codex tool-mapping table and detailed team-mode rules that could equally be pushed into the reference bundle.

4 / 5

Total

15

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description answers both what and when with an explicit trigger clause, but it is held back by abstract jargon ('ultrawork', 'evidence-bound', 'systematic') and a thin set of natural trigger terms. Adding concrete capability verbs and the synonyms already used in the body would lift specificity and trigger quality.

Suggestions

Replace abstract phrasing ('systematic, evidence-bound ultrawork steps') with concrete capability verbs, e.g. 'Registers goals, delegates work to subagents, captures per-criterion evidence at a git tree, and merges verified units atomically.'

Broaden the 'Use when' clause with the natural variations the body already lists: 'Use when the user asks for ulw-loop, ulw, a goal loop, durable or checkpointed long-running execution, or evidence-led work.'

State a distinguishing qualifier so it cannot fire for generic goal/planning skills, e.g. 'for long-running, checkpointed delivery with per-criterion evidence' rather than just 'goal loop'.

DimensionReasoningScore

Specificity

The description names its domain ('goal-like loop') and 1-2 actions ('decomposes work', 'evidence-bound ultrawork steps'), but 'systematic' and 'ultrawork' are abstract jargon rather than the several concrete actions needed for a 4.

3 / 5

Completeness

Both 'what' and an explicit 'Use when...' clause are present; the 'when' carries concrete trigger phrases, but the 'what' leans on the jargon term 'ultrawork' rather than the plain concrete capability list required for a 5.

4 / 5

Trigger Term Quality

'goal loop' and 'durable, checkpointed execution' are relevant natural keywords, but common variations the body itself relies on ('ulw', 'ulw-loop', 'manual QA', 'long-running delivery') are missing from the description.

3 / 5

Distinctiveness Conflict Risk

'goal loop or durable, checkpointed execution' carves a clear niche, but 'goal loop' could still overlap with generic planning or goal-management skills, keeping it short of the minimal-conflict 5 anchor.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
code-yeongyu/oh-my-openagent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.