CtrlK
BlogDocsLog inGet started
Tessl Logo

autoresearch

Stateful single-mission improvement loop with strict evaluator contract, markdown decision logs, and max-runtime stop behavior

53

Quality

60%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/autoresearch/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is lean, well-structured, and clear about the single-mission contract and artifact layout. Its main weaknesses are the absence of executable/validated workflow steps and a missing error-recovery feedback loop for the iterative evaluation cycle.

Suggestions

Add an explicit validation checkpoint after persisting evaluation JSON (e.g. confirm the file is well-formed and contains the required 'pass' field) before appending the decision log.

Provide at least one concrete runnable example of invoking the evaluator and writing an evaluation JSON entry, rather than only describing it abstractly.

Include a short feedback-loop note for the non-passing case (how to decide the next experiment from a failed evaluation) so the iterate->evaluate->persist cycle is actionable.

DimensionReasoningScore

Conciseness

The body is efficient and assumes Claude's competence, using compact tagged sections that avoid explaining concepts Claude already knows; only minor phrasing could be tightened further.

4 / 5

Actionability

Guidance is concrete in artifact shape (directory layout, JSON fields) but the workflow steps are directives rather than executable commands, and the evaluator is referenced abstractly without a runnable example.

3 / 5

Workflow Clarity

The iteration loop is clearly sequenced, but it involves batch/iterative destructive operations with no explicit validation checkpoint (e.g. verifying evaluation JSON is valid/persisted) and no error-recovery feedback loop, which caps clarity at 3.

3 / 5

Progressive Disclosure

Content is well-organized into focused tagged sections with a compact inline artifact shape and clear separation of concerns; no bundle files exist so references are minimal and appropriately one-level, with only minor organization gaps.

4 / 5

Total

14

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and technically descriptive about the loop's mechanics, but it lacks an explicit 'when to use' trigger and relies on internal jargon rather than natural user keywords. It is distinguishable but not user-trigger-friendly.

Suggestions

Add an explicit 'Use when...' clause naming natural user phrases (e.g. iteratively improving a mission against an evaluator, bounded improvement runs).

Replace or augment internal terms ('evaluator contract', 'max-runtime stop behavior') with user-facing synonyms so natural-language triggering works.

Lead with the concrete action (iterate/run experiments against an evaluator until a runtime ceiling) before the implementation mechanisms.

DimensionReasoningScore

Specificity

Names concrete mechanics ('stateful single-mission improvement loop', 'strict evaluator contract', 'markdown decision logs', 'max-runtime stop behavior'), listing several specific actions rather than vague domain references, though it omits concrete actions a user would invoke.

4 / 5

Completeness

The 'what' is clear (bounded evaluator-driven iterative improvement with durable logs), but there is no 'Use when...' clause or explicit trigger guidance, which per the guidelines caps completeness at 3.

3 / 5

Trigger Term Quality

Contains relevant skill-specific terms ('autoresearch', 'evaluator', 'max-runtime', 'cron') but these are technical/internal labels rather than the natural phrases a user would say when they need this skill, and no synonyms or user-facing triggers are given.

3 / 5

Distinctiveness Conflict Risk

The combination of 'stateful single-mission improvement loop' with a 'strict evaluator contract' is a fairly distinct niche with limited overlap risk, though adjacent automation/research skills could partially overlap.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
Yeachan-Heo/oh-my-claudecode
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.