CtrlK
BlogDocsLog inGet started
Tessl Logo

autoresearch

Autonomous Goal-directed Iteration. Apply Karpathy's autoresearch principles to ANY task. Loops autonomously — modify, verify, keep/discard, repeat. 9 subcommands: plan, debug, fix, security, ship, scenario, predict, learn.

51

Quality

56%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/autoresearch/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

62%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-engineered operational document — the loop protocol, setup gating, and validation/retry logic are exemplary — but it violates its own progressive-disclosure design by inlining six full subcommand manuals' worth of flags, usage, and metrics that duplicate 11 dedicated reference files. Cutting the body to the trigger map, setup gate, core loop, and one-line-per-subcommand pointers would roughly halve its token cost with no loss of capability.

Suggestions

Move each subcommand's flag tables, usage blocks, and composite-metric formulas into the corresponding references/<subcommand>-workflow.md and replace them with a 2–3 line summary plus the 'Load:' pointer — this is the single biggest win for both conciseness and progressive disclosure.

Deduplicate the interactive setup material: the 'Interactive Setup Gate' table and the 'Setup Phase' section repeat the same per-command question requirements; keep one canonical table and reference it.

Resolve the iteration-count syntax inconsistency — the body teaches 'Iterations: N' inline config while the CI/CD row uses '--iterations N' with no usage example — and define the variables used in composite metric formulas.

Trim the 'When to Activate' list to one representative trigger phrase per subcommand; the current ~20-phrase list restates the subcommand table at length.

DimensionReasoningScore

Conciseness

The 601-line body is noticeably verbose: the Interactive Setup Gate table is restated almost in full in the 'Setup Phase' section, the 'When to Activate' list enumerates ~20 trigger phrases per subcommand that mostly restate the subcommand table, and each of the six subcommand sections carries flag tables, long usage blocks, and metric formulas that duplicate the dedicated reference files. Not 3 because the padding is pervasive across multiple sections rather than isolated; not 1 because it does not explain concepts Claude already knows — nearly all of it is operational content, just inlined in the wrong place.

2 / 5

Actionability

Guidance is highly concrete and executable: exact invocations with flags and inline config ('/autoresearch:security --diff --fix --fail-on critical'), the dry-run-the-verify-command-before-accepting gate, explicit decision rules with retry caps, output file names, and batched AskUserQuestion tables with option lists. Not 5 because of the unresolved duality between the inline 'Iterations: 25' syntax and the '--iterations N' flag that appears only in the CI/CD table with no usage example, plus metric formulas whose variables (e.g. 'min(findings, 20)') are never defined in the body.

4 / 5

Workflow Clarity

The Loop is an explicitly sequenced 9-step procedure with validation at every checkpoint: baseline verification as iteration #0, dry-run of the verify command before launch, guard execution each iteration, mechanical pass/fail decision rules with rollback ('git revert, not git reset --hard'), capped retries (2 for guard conflicts, 3 for crashes), and mandatory logging. Destructive/batch risk is well-covered by these feedback loops, so the workflow-clarity cap does not apply. Not 4 because validation steps are explicit and interleaved rather than implicit or partly missing.

5 / 5

Progressive Disclosure

Structure is real and the navigation works — each subcommand section opens with 'Load: references/<subcommand>-workflow.md for full protocol', and all 11 referenced files exist with no second-level references (verified one level deep). However, roughly 300 lines of per-subcommand flag tables, usage blocks, and composite-metric formulas are inlined in SKILL.md when that detail self-evidently belongs in the already-existing reference files, matching the anchor of overview content that should be separate sitting inline. Not 4 because the duplication is substantial rather than a minor organization gap; not 2 because references are prominent and clearly signaled, not buried.

3 / 5

Total

14

/

20

Passed

Description

51%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description communicates the core mechanism and command inventory clearly, but lacks any explicit 'when to use' guidance and over-claims applicability to 'ANY task', creating conflict risk with narrower dedicated skills. It also contains a factual inconsistency ('9 subcommands' but only 8 listed), which undermines precision. Overall it sits just above the midpoint: concrete but not trigger-ready.

Suggestions

Add an explicit 'Use when...' clause with concrete trigger phrases (e.g. 'Use when the user says iterate until done, run overnight, keep improving, hunt bugs, threat model, or ship it'), which would raise completeness above the cap of 3.

Fix the count mismatch ('9 subcommands' vs the 8 actually listed: plan, debug, fix, security, ship, scenario, predict, learn).

Narrow the 'ANY task' over-claim to the actual domain (tasks with a mechanically verifiable metric) and state what distinguishes it from built-in review, security, and docs skills to reduce conflict risk.

Briefly say what the subcommands do (e.g. 'security: STRIDE/OWASP audit; ship: pre-ship checklist workflow') so the command list doubles as a capability list.

DimensionReasoningScore

Specificity

Concrete loop mechanics ('Loops autonomously — modify, verify, keep/discard, repeat') and eight named subcommands are specific actions, but the description never states what each subcommand does or what kinds of work it improves, leaving coverage gaps. Not 5 because the actions are generic loop verbs rather than a comprehensive list of concrete capabilities; not 3 because it does name several specific actions and a full command inventory beyond just naming the domain.

4 / 5

Completeness

The 'what' is clear — an autonomous modify→verify→keep/discard loop derived from Karpathy's autoresearch — but there is no 'Use when...' or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines. 'Apply ... to ANY task' only weakly implies a when, and is too vague to count as an explicit trigger. Not 4 because any 'when' clause is absent rather than merely under-specified.

3 / 5

Trigger Term Quality

Subcommand names ('debug', 'fix', 'security', 'ship', 'plan') double as natural user keywords, and 'Loops autonomously' is a plausible phrase, but there are no situational trigger phrases ('iterate until done', 'run overnight', 'threat model', 'ship it') or synonyms beyond the bare command list. Not 4 because common variations users would actually say are absent; not 2 because the listed terms are relevant and natural rather than purely technical jargon.

3 / 5

Distinctiveness Conflict Risk

'Apply ... to ANY task' claims universal applicability, creating high overlap risk with virtually every other skill, and the subcommands (plan, debug, fix, security, ship, learn) collide directly with dedicated planning, debugging, security-audit, and documentation skills. Not 3 because the breadth is explicitly unlimited rather than merely 'somewhat specific'; not 1 because the autoresearch framing and named commands give it some identity beyond pure generic language.

2 / 5

Total

12

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (602 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
OpenLAIR/dr-claw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.