CtrlK
BlogDocsLog inGet started
Tessl Logo

autoresearch-create

Set up and run an autonomous experiment loop for any optimization target. Gathers what to optimize, then starts the loop immediately. Use when asked to "run autoresearch", "optimize X in a loop", "set up autoresearch for X", or "start experiments".

70

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized, highly actionable skill body with concrete templates and validation mechanisms for its batch/destructive loop. It is slightly verbose in places and remains monolithic where splitting some reference material could aid discovery.

Suggestions

Tighten explanatory prose (e.g. the .auto rationale and measure.sh preamble) to trim tokens without losing the actionable core.

Consider extracting the .auto/prompt.md template or .auto/config.json field reference into a separate reference file so SKILL.md reads as a leaner overview with one-level-deep links.

Add an explicit validate-then-proceed checklist to the Setup sequence (e.g. "baseline must pass checks.sh before looping") to consolidate the distributed checkpoints.

DimensionReasoningScore

Conciseness

Mostly lean and actionable with a few explanatory asides (e.g. "This keeps everything in one place", "every second is multiplied by hundreds of runs") that could be trimmed, stopping just short of the fully efficient 5-anchor.

4 / 5

Actionability

Provides copy-paste-ready bash, JSON, and markdown template examples plus concrete tool calls and file paths (.auto/prompt.md, .auto/measure.sh, init_experiment) covering the common cases.

5 / 5

Workflow Clarity

Setup is a clear 5-step sequence with validation present (checks.sh, confidence score, "cannot keep when checks failed"), but checkpoints are distributed rather than a single explicit validate-then-proceed checklist.

4 / 5

Progressive Disclosure

Well-organized into clear sections with no nested or broken references, but as a >50-line single-file doc with no external bundle references it lacks the one-level-deep reference split that defines the 5-anchor.

4 / 5

Total

17

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description that concretely states capabilities and provides explicit, natural trigger phrases. It clearly answers both what the skill does and when to invoke it with minimal conflict risk.

DimensionReasoningScore

Specificity

Lists several concrete actions — "Set up and run an autonomous experiment loop", "Gathers what to optimize", "starts the loop immediately" — with only minor granularity gaps versus the comprehensive 5-anchor.

4 / 5

Completeness

Explicitly answers both what ("Set up and run an autonomous experiment loop for any optimization target") and when ("Use when asked to ...") with concrete trigger phrases, matching the 5-anchor.

5 / 5

Trigger Term Quality

Provides multiple natural trigger phrases ("run autoresearch", "optimize X in a loop", "set up autoresearch for X", "start experiments") users would actually say, missing only minor synonym variation to reach 5.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche ("autoresearch" / autonomous experiment loop) with distinct triggers and minimal overlap risk with other skills.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
davebcn87/pi-autoresearch
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.