CtrlK
BlogDocsLog inGet started
Tessl Logo

weights-and-biases

W&B: log ML experiments, sweeps, model registry, dashboards.

52

Quality

60%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/mlops/evaluation/weights-and-biases/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill is rich with executable, framework-spanning examples and points to genuine reference files, but it over-explains basics, carries marketing and pricing material, and inlines content that duplicates its own bundle files. Tightening and pushing detail into references would materially raise conciseness and progressive disclosure.

Suggestions

Remove basic-concept glosses ('Project: Collection of related experiments') and the marketing/Pricing sections that Claude does not need to act.

Move the full sweeps, artifacts, and integrations code blocks into the existing references/*.md files and keep SKILL.md as a concise overview pointing to them.

Collapse the three repeated sweep-config examples into one with the method variants noted inline.

DimensionReasoningScore

Conciseness

The body glosses concepts Claude already knows ('Project: Collection of related experiments'), carries marketing ('200,000+ users', '10.5k+ stars'), a Pricing section, and repeats sweep config three times — noticeably verbose with several padded sections.

2 / 5

Actionability

Substantial concrete, copy-paste-ready Python across tracking, PyTorch, sweeps, artifacts, and integrations; minor gaps from undefined placeholders like train_epoch()/validate() keep it just below fully executable.

4 / 5

Workflow Clarity

Content is organized by feature area rather than as a sequenced workflow, and validation/checkpoint steps are implicit; the offline-mode and best-practices tips are not framed as validate-then-fix loops.

3 / 5

Progressive Disclosure

Three real reference files (sweeps.md, artifacts.md, integrations.md) are clearly signaled one level deep in 'See Also', but the body inlines full sweeps/artifacts/integrations code that largely duplicates those dedicated files.

3 / 5

Total

12

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and names W&B's core capabilities concisely, but it omits any explicit trigger ('Use when...') clause, which is the most consequential gap. Adding when-to-use guidance and a few synonyms would lift it toward the top anchors.

Suggestions

Add a 'Use when...' clause naming concrete triggers, e.g. 'Use when logging ML experiments, running hyperparameter sweeps, or managing a model registry.'

Include the full name and library synonym — 'Weights & Biases (wandb)' — so users who say either form match the skill.

Drop or move the marketing-flavored breadth (four nouns) in favor of trigger phrases a user would actually say.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'log ML experiments, sweeps, model registry, dashboards' — giving comprehensive coverage of the tool's capabilities within a tight phrase.

5 / 5

Completeness

A clear 'what' is stated but there is no 'Use when...' or equivalent trigger guidance, which per the judging guidelines caps completeness at 3.

3 / 5

Trigger Term Quality

Natural practitioner terms ('ML experiments', 'sweeps', 'model registry', 'dashboards') are present, but synonyms like 'Weights & Biases', 'wandb', 'experiment tracking', and 'hyperparameter tuning' are missing.

4 / 5

Distinctiveness Conflict Risk

W&B is a specific named tool with mostly distinct triggers; minor overlap risk with adjacent MLOps skills via generic terms like 'dashboards' and 'ML experiments'.

4 / 5

Total

16

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (599 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.