CtrlK
BlogDocsLog inGet started
Tessl Logo

autoresearch

Run bounded automated experiment iterations by recording baselines, applying hypothesis patches, comparing metrics, protecting regression guards, and deciding keep, discard, rollback, or block. Use when automated research is requested or a repo/skill needs evidence-backed research, metric tracking, or safe optimisation loops.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Failed to scan

The risk profile of this skill

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Well-structured content with executable examples, explicit validation feedback loops, and clean one-level-deep references. Minor conciseness and actionability polish would reach the top anchors.

DimensionReasoningScore

Conciseness

Lean, dense prose with tight headers and economical bullets, assuming Claude's competence; only minor phrases could be trimmed, so it sits just below fully lean.

4 / 5

Actionability

Provides concrete ledger YAML and executable shell commands (uv run train.py, ./bin/ask evals run), with minor gaps such as the bare 'apply_patch' placeholder and some procedural-only guidance.

4 / 5

Workflow Clarity

A 9-step workflow with explicit validation checkpoints (baseline first, guard checks, held-out checks, fail-fast) and keep/discard/block feedback loops for batch and destructive operations.

5 / 5

Progressive Disclosure

A dedicated Progressive Disclosure section points one level deep to real bundle files (autoresearch-project.md, contract.yaml, evals.yaml, task-profile.json), keeping the body an overview with clear navigation.

5 / 5

Total

18

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description with explicit trigger guidance and concrete capabilities. Slight headroom on trigger-term variety and distinctiveness versus broader research skills.

DimensionReasoningScore

Specificity

Lists five concrete actions (recording baselines, applying hypothesis patches, comparing metrics, protecting regression guards, deciding keep/discard/rollback/block), matching the comprehensive-coverage anchor.

5 / 5

Completeness

Explicitly states what the skill does and includes a concrete 'Use when automated research is requested or a repo/skill needs...' trigger clause, satisfying both what and when.

5 / 5

Trigger Term Quality

Natural terms like 'automated research', 'evidence-backed research', 'metric tracking', and 'safe optimisation loops' are present, but coverage leans technical and lacks synonym/extension variants, falling just below comprehensive.

4 / 5

Distinctiveness Conflict Risk

The bounded-experiment-loop niche with regression guards is mostly distinct, but 'research', 'metric tracking', and 'optimisation loops' could overlap with general experimentation skills.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

Total

15

/

16

Passed

Repository
jscraik/Agent-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.