CtrlK
BlogDocsLog inGet started
Tessl Logo

run-experiment

Deploy and run ML experiments on local, remote, Vast.ai, or Modal serverless GPU. Use when user says "run experiment", "deploy to server", "跑实验", or needs to launch training jobs.

61

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

High

Do not use without reviewing

Fix and improve this skill with Tessl

tessl review fix ./skills/run-experiment/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced operations runbook with concrete commands and real verification checkpoints for each environment. Its weaknesses are token efficiency (inline W&B boilerplate and repeated Modal/Vast guidance) and progressive disclosure — everything lives in one ~310-line SKILL.md with no bundle files, where per-provider guides and the CLAUDE.md example would be better split out.

Suggestions

Cut the W&B section to the decision points (when to inject logging, which metric names, where keys come from) and drop the wandb.init/log/finish boilerplate Claude already knows.

Split per-provider detail (Vast.ai lifecycle, Modal launcher recipe, the CLAUDE.md example) into references/ files with one-level-deep pointers, keeping SKILL.md as the routing overview.

Add an explicit pre-destroy checkpoint in Step 7 that verifies the results rsync/scp succeeded (e.g., check file count or size) before running 'vastai destroy instance'.

DimensionReasoningScore

Conciseness

The bulk is lean executable commands, but the W&B section (~40 lines) teaches boilerplate Claude already knows ("import wandb / wandb.init / wandb.log / wandb.finish" usage patterns), and Modal/Vast guidance repeats across Step 1, Step 4, Key Rules, and the trailing setup notes. This is more than the 'minor instances' of anchor 4 but the body is far from the padded explanatory style of anchor 2.

3 / 5

Actionability

Highly concrete guidance throughout: exact rsync include/exclude flag lists, nvidia-smi query commands, full screen launch strings, and vastai destroy invocations, with placeholders sourced from documented files (vast-instances.json, CLAUDE.md). Anchor 5 is not reached because most blocks are templates requiring substitution and the W&B snippet contains pseudocode ("{...hyperparams...}") rather than copy-paste-ready code.

4 / 5

Workflow Clarity

A clear 7-step sequence with per-environment branches, explicit skip conditions, and real checkpoints (Step 5 'screen -ls' / 'modal app list' verification, the 'memory.used < 500 MiB' free-GPU test, 'wandb status' login check). The gap that keeps it at anchor 4: the destructive 'vastai destroy instance' step has a completion trigger but no verification that the preceding results rsync/scp succeeded before destroying the instance.

4 / 5

Progressive Disclosure

No bundle files exist (no references/, scripts/, or assets/), and the ~310-line body is monolithic: content that clearly belongs in separate reference files — the W&B integration guide, the ~30-line CLAUDE.md example, per-provider deep-dives — is inlined. Section headers are well organized and the one external pointer (../shared-references/compute-env-contract.md) is clearly signaled, matching anchor 3 rather than anchor 2's minimal structure or anchor 4's appropriately split content.

3 / 5

Total

14

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete capability statement across four named environments, and an explicit 'Use when' clause with natural trigger phrases including a Chinese synonym. The only weaknesses are incomplete synonym coverage and minor trigger overlap with the sibling Modal/Vast skills it references.

Suggestions

Add trigger variations like "train a model" or "run training" to broaden natural keyword coverage.

Briefly surface the full lifecycle (rent/provision → sync → run → collect → destroy) to sharpen what-distinctiveness versus the /serverless-modal and /vast-gpu skills it delegates to.

DimensionReasoningScore

Specificity

"Deploy and run ML experiments on local, remote, Vast.ai, or Modal serverless GPU" names two concrete actions and enumerates four specific deployment targets with the serverless qualifier. It falls between anchor 3 (1-2 concrete actions) and anchor 5 (comprehensive multi-action coverage) — the target enumeration adds specificity beyond anchor 3, but the full lifecycle verbs (provision, sync, monitor, destroy) are absent, so it is not comprehensive.

4 / 5

Completeness

Explicitly answers both questions: the "what" is "Deploy and run ML experiments on local, remote, Vast.ai, or Modal serverless GPU", and the "when" is a concrete "Use when user says..." clause listing trigger phrases. This matches anchor 5 exactly; anchor 4 would require the when-clause to be less explicit or specific.

5 / 5

Trigger Term Quality

Includes natural trigger phrases users would actually say — "run experiment", "deploy to server", "跑实验", "launch training jobs" — with a Chinese-language synonym. Coverage is good but misses common variations such as "train a model", "run training", or "start a job", so it does not reach the comprehensive synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

ML experiment deployment is a clear niche with distinct triggers ("run experiment", "deploy to server"), but the description explicitly names Modal serverless and Vast.ai GPU territory that overlaps with the sibling /serverless-modal and /vast-gpu skills it delegates to. Mostly distinct with minor overlap risk against closely related skills — anchor 4 rather than anchor 5's minimal conflict risk.

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.