CtrlK
BlogDocsLog inGet started
Tessl Logo

monitor-experiment

Monitor running experiments, check progress, collect results. Use when user says "check results", "is it done", "monitor", or wants experiment output.

67

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable monitoring skill with executable snippets and clear environment-specific branches. Weakest on validation feedback loops and a dangling external reference.

Suggestions

Add an explicit validate->fix->retry checkpoint (e.g., after collecting output, confirm JSON parses and runs are non-empty before summarizing) to strengthen workflow_clarity.

Fix or remove the broken reference to ../shared-references/external-cadence.md, or move the inlined W&B/Modal/Vast detail into bundle reference files to improve progressive_disclosure.

Trim the external-cadence blockquote and the "What to extract" explanatory prose to tighten conciseness while keeping the actionable bullets.

DimensionReasoningScore

Conciseness

Mostly efficient with executable snippets and terse bullets, but the external-cadence blockquote and some explanatory prose ("is it converging? diverging?") add tokens that could be trimmed.

4 / 5

Actionability

Provides copy-paste-ready bash and python snippets covering screen capture, JSON results, W&B metrics, Modal, and a comparison-table template across the common cases.

5 / 5

Workflow Clarity

A clear numbered Step 1–6 sequence with conditional gating per environment, but validation checkpoints and error-recovery feedback loops are only minimally present.

4 / 5

Progressive Disclosure

Well-organized into steps with one clearly-signaled one-level reference (external-cadence.md), but that referenced file does not exist in the repo and no bundle files are provided to offload the inlined W&B/Modal/Vast detail.

4 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that clearly states capabilities and explicit trigger phrases. It is held back from perfection only by slightly generic action verbs and minor trigger overlap risk.

DimensionReasoningScore

Specificity

Lists several concrete actions ("Monitor running experiments, check progress, collect results") but the verbs are somewhat generic and not the comprehensive multi-action coverage of the top anchor.

4 / 5

Completeness

Explicitly answers both what ("Monitor running experiments, check progress, collect results") and when with concrete trigger phrases ("Use when user says...").

5 / 5

Trigger Term Quality

Includes natural phrases users would actually say ("check results", "is it done", "monitor") plus "wants experiment output", with good coverage but a few natural synonyms missing.

4 / 5

Distinctiveness Conflict Risk

The experiment-monitoring niche is mostly distinct, but bare terms like "check results" and "monitor" carry minor overlap risk with more generic skills.

4 / 5

Total

17

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 suspicious

Warning

Total

13

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.