CtrlK
BlogDocsLog inGet started
Tessl Logo

run-experiment

Deploy and run ML experiments on local or remote GPU servers. Use when user says "run experiment", "deploy to server", "跑实验", or needs to launch training jobs.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

High

Do not use without reviewing

Fix and improve this skill with Tessl

tessl review fix ./skills/skills-codex/run-experiment/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-sequenced, highly actionable runbook with real validation checkpoints around risky operations. Its weaknesses are inlined detail that belongs in reference files and a few verbose sections that could be trimmed.

Suggestions

Move the W&B integration details and the AGENTS.md example template into files under references/ and link them from the body, keeping SKILL.md a lean overview.

Tighten the dense 'Environment contract' paragraph into a short bullet list or move it into the referenced contract file.

Give Step 6 (Feishu notification) a concrete command or payload example instead of only describing what to send.

DimensionReasoningScore

Conciseness

Mostly efficient with terse bullets and command blocks, but the dense 'Environment contract' paragraph, the multi-line W&B python template, and the full AGENTS.md example could be tightened or offloaded. Fits anchor 3 ('mostly efficient but includes some unnecessary explanation or could be tightened'); not 4 given the wordy env-contract prose.

3 / 5

Actionability

Concrete, near-executable commands throughout (nvidia-smi query, rsync filter rules, the full screen + conda-hook launch line, 'modal run'). Minor gaps: the Feishu step gives no concrete send command, and 'config={...hyperparams...}' is placeholder pseudocode. Anchor 4 ('concrete code or commands with minor gaps'); not 5 because of those gaps.

4 / 5

Workflow Clarity

A clear 7-step sequence with explicit validation checkpoints: pre-flight GPU check, a dedicated 'Verify Launch' step, and a guarded auto-destroy flow ('Verify the target process has exited', 'If any artifact copy fails, do not destroy') plus error-recovery guidance for Modal/Vast failures. Destructive-operation validation is present, so the cap does not apply; matches anchor 5.

5 / 5

Progressive Disclosure

Header structure is good, but no bundle files exist and the ~50-line W&B section and AGENTS.md template are inlined rather than split into references. The sole external reference ('../shared-references/compute-env-contract.md') is a buried parenthetical rather than clearly signaled navigation. Fits anchor 3 ('content that should be separate is inline; references not clearly signaled').

3 / 5

Total

15

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use when...' clause and natural trigger phrases including a multilingual synonym. The only weakness is thin action coverage — it says 'deploy and run' without hinting at GPU checks, session isolation, or environment variety.

Suggestions

Enumerate one or two more concrete capabilities in the description (e.g., GPU availability checks, screen/background session management) to lift specificity.

Add trigger variants users commonly say, such as 'train model' or 'start training', to broaden trigger-term coverage.

DimensionReasoningScore

Specificity

Names the domain ('ML experiments', 'local or remote GPU servers') and two concrete actions ('Deploy and run'), but stops short of listing several specific capabilities. Matches anchor 3 (domain plus 1-2 concrete actions, not comprehensive); not 4 because it does not enumerate multiple specific actions.

3 / 5

Completeness

Explicitly answers both what ('Deploy and run ML experiments on local or remote GPU servers') and when ('Use when user says...') with concrete quoted trigger phrases, closely mirroring the anchor-5 example. Not 4 since the 'when' clause is already explicit and specific.

5 / 5

Trigger Term Quality

Natural trigger phrases are present ('run experiment', 'deploy to server', 'launch training jobs', plus the Chinese synonym '跑实验'). Not 5: common variations like 'train model', 'start training', or 'launch a run' are missing; not 3 because coverage clearly exceeds 'some relevant keywords'.

4 / 5

Distinctiveness Conflict Risk

A clear niche (ML experiment deployment on GPU servers) with distinct triggers; unlikely to fire for unrelated skills. Matches the anchor-5 'clear niche with distinct triggers; minimal conflict risk'.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.