CtrlK
BlogDocsLog inGet started
Tessl Logo

autolab-managed-experiment

Run one Autolab benchmark experiment safely on Hugging Face Jobs. Use when a planner, reviewer, or experiment worker is preparing, auditing, launching, or reviewing a single train.py hypothesis against the current local promoted master.

91

1.94x
Quality

90%

Does it follow best practices?

Impact

99%

1.94x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a tight, executable, well-sequenced instruction skill with explicit validation checkpoints and stop-condition guardrails. Its only meaningful gap is leaving the JOB_ID and comment placeholders unillustrated rather than giving a concrete worked example.

Suggestions

Add a brief worked example showing a concrete JOB_ID and --comment value (e.g. `hf_job.py logs 42 --follow ...` and `submit_patch.py --comment "lr=3e-4 val_bpb=0.42"`) so the logs/submit steps are fully copy-paste ready.

Make the metric-parsing expectation explicit (e.g. note what `parse_metric.py` emits and where to record it) so the record step is unambiguous.

Optionally cross-reference the 'Fast Checks' preflight --json output as the input to the 'stop and inspect the diff' guardrail, tying the validation loop together.

DimensionReasoningScore

Conciseness

Lean and efficient: the body is almost entirely executable commands and tightly scoped guardrails, with no over-explanation of concepts Claude already knows. Every section earns its place, including the non-obvious git-main guardrail.

5 / 5

Actionability

Provides concrete, executable `uv run scripts/...` commands with concrete output paths and the preflight/launch/logs/parse sequence. It falls short of 5 because placeholders like `<JOB_ID>` and `--comment "..."` are left unfilled rather than shown with a worked example.

4 / 5

Workflow Clarity

A clearly numbered 7-step sequence with an explicit preflight validation checkpoint before launch and explicit stop-and-inspect / stop-and-rewrite guardrails as feedback loops for risky operations, matching the anchor with validation steps and error-recovery guidance.

5 / 5

Progressive Disclosure

Under 50 lines with no external bundle files, organized into well-signaled sections (Workflow, Guardrails, Fast Checks) where the Fast Checks clearly separate programmatic/auxiliary commands from the main flow; the simple-skill exception applies.

5 / 5

Total

19

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, complete, and well-scoped to a niche domain, with an explicit 'Use when' clause and concrete trigger roles. It is strong overall, with the only minor weakness being somewhat domain-jargon-heavy trigger terms rather than broadly natural synonyms.

DimensionReasoningScore

Specificity

Names the domain ('Autolab benchmark experiment safely on Hugging Face Jobs') and several concrete actions ('preparing, auditing, launching, or reviewing a single train.py hypothesis'), with only minor coverage gaps. Not a 5 because the action list, while specific, is narrow and does not enumerate the full set of operations a single experiment entails.

4 / 5

Completeness

Explicitly answers both 'what' ('Run one Autolab benchmark experiment safely on Hugging Face Jobs') and 'when' ('Use when a planner, reviewer, or experiment worker is preparing, auditing, launching, or reviewing a single train.py hypothesis...') with concrete trigger phrases, matching the 5-anchor example pattern.

5 / 5

Trigger Term Quality

Includes natural terms users would say ('Autolab benchmark experiment', 'Hugging Face Jobs', 'train.py hypothesis') plus role-based triggers ('planner, reviewer, or experiment worker'). Good keyword coverage, but it is jargon-heavy and lacks common synonyms or file-extension variants that would push it to 5.

4 / 5

Distinctiveness Conflict Risk

A clear niche scoped to a single Autolab train.py hypothesis on Hugging Face Jobs with highly specific terminology ('local promoted master'); minimal overlap risk with other skills.

5 / 5

Total

18

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 9 missing

Warning

Total

12

/

16

Passed

Repository
huggingface/context-course
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.