CtrlK
BlogDocsLog inGet started
Tessl Logo

autoresearch

Stateful validator-gated research loop with native-hook persistence

55

Quality

69%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/oh-my-codex/skills/autoresearch/SKILL.md

The canonical home for this skill is autoresearch in Yeachan-Heo/oh-my-codex

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a compact, well-structured orchestration contract: it states concrete state files, artifact schemas, and invocation syntax, and its validation-gated completion loop is exceptionally explicit. Minor duplication between the Core contract and Migration note and a couple of elided specifics keep it just short of top marks.

DimensionReasoningScore

Conciseness

The body is tight and assumes competence: no concept explanations, and concrete specifics like the state-file fields and passing-artifact JSON examples. It falls short of lean-perfection only through minor duplication — 'Direct CLI launch is gone' in Core contract item 4 is restated as 'No direct CLI launch. No tmux split-pane launch.' in the Migration note, and 'hard-deprecated' appears twice — matching the 'efficient; minor instances that could be trimmed' anchor.

4 / 5

Actionability

Concrete, executable guidance throughout: exact state path (`.omx/state/.../autoresearch-state.json`) with an enumerated field list, full example JSON artifacts with concrete paths, exact invocation syntax (`$deep-interview --autoresearch`), and a numbered flow naming the files to materialize. It stops short of copy-paste-ready because the `...` path segment is elided and no example `mission_validator_command` is given — 'mostly executable guidance with minor gaps'.

4 / 5

Workflow Clarity

The 'Recommended flow' is a clear 5-step sequence whose centerpiece is an explicit validation gate: 'Let stop-hook / auto-nudge continue until the completion artifact satisfies the chosen validation mode' plus the rule that the loop 'does not stop because the model says "done", because a stop hook fired once, or because several turns were no-ops'. Both validation modes have concrete passing-artifact examples defining acceptance criteria, and the loop-until-validated design is itself the validate→continue feedback loop. Validation here is not merely 'mostly present' (anchor 4) — it is the explicit core contract.

5 / 5

Progressive Disclosure

No bundle files exist, and the ~65-line body is organized into clear, well-labeled sections (Boundary, Use when, Do not use when, Core contract, Completion artifact contract, Recommended flow, Migration note) that are easy to navigate. It scores 4 rather than 5 because the file runs a bit denser than a lean overview — the inline JSON artifact examples push it past the under-50-line self-contained case the scoring notes exempt — though nothing currently warrants splitting.

4 / 5

Total

17

/

20

Passed

Description

41%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description conveys a distinct internal mechanism but reads as engineering jargon: it says what the skill is at a technical level while providing no 'use when' guidance and almost no natural keywords a user would actually say. It has a clear niche identity but weak user-facing discoverability.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user wants a persistent research loop that keeps running until an explicit validator passes.'

Replace jargon modifiers with concrete plain-language actions, e.g. 'Runs repeated research passes and only stops once a validator approves the output; persists loop state across sessions.'

Include natural synonyms and phrasings users would say — 'automated research', 'research loop', 'keep researching until validated' — so the description matches how the need is actually voiced.

DimensionReasoningScore

Specificity

The description names the domain ("research loop") but expresses capabilities only as jargon modifiers — "Stateful", "validator-gated", "native-hook persistence" — with no concrete actions or verbs. It sits at the 'names the domain but actions are minimal or generic' anchor: more substantive than pure abstraction (not 1), but unlike the score-3 anchor it lists no concrete actions like 'extracts text' or 'fills forms'.

2 / 5

Completeness

The 'what' is stated (a validator-gated persistent research loop) but the 'when' is entirely absent — no 'Use when...' clause or equivalent trigger guidance, which caps this dimension at 3 per the judging guidelines. It is not 2 because the 'what' is specific rather than vague, and not 4 because the 'when' is missing outright rather than merely under-specified.

3 / 5

Trigger Term Quality

"research" is the single quasi-natural keyword; everything else ("validator-gated", "stateful", "native-hook persistence") is technical jargon no user would say, and there are no natural trigger phrases or synonyms. This matches the score-2 anchor ('one or two generic keywords; missing the natural phrases users say') better than score 1 only because 'research' is a term users genuinely utter.

2 / 5

Distinctiveness Conflict Risk

The niche is fairly distinct — "validator-gated research loop with native-hook persistence" is unlikely to trigger for unrelated skills — with only minor overlap risk against general research or iteration skills. It falls short of 5 because no explicit trigger phrases sharpen the boundary against those closely related research skills.

4 / 5

Total

11

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Yeachan-Heo/oh-my-codex
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.