CtrlK
BlogDocsLog inGet started
Tessl Logo

weights-and-biases

W&B: log ML experiments, sweeps, model registry, dashboards.

60

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/mlops/evaluation/weights-and-biases/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with executable examples, but it is verbose and duplicates content already available in the reference files, with no validation checkpoints for batch operations like sweeps.

Suggestions

Trim marketing stats, the pricing section, and redundant sweep-strategy examples to reduce token cost; keep one sweep config with inline comments for the method variants.

Move the detailed sweeps, artifacts, and integrations sections into their reference files and replace the inline content with brief one-line pointers, leaving SKILL.md as an overview.

Add explicit validation/verification steps for sweeps and artifact logging (e.g. confirm run URL is live, verify artifact download checksum) to raise workflow clarity.

DimensionReasoningScore

Conciseness

The ~588-line body is mostly useful executable code, but it is padded with marketing stats ('200,000+ ML practitioners', 'GitHub Stars: 10.5k+'), a pricing section, and redundant sweep-strategy blocks that could be tightened.

2 / 3

Actionability

Abundant concrete, executable, copy-paste-ready examples for init, logging, sweeps, artifacts, and framework integrations meet the 'fully executable code' anchor.

3 / 3

Workflow Clarity

An implicit install→init→log→finish flow exists in Quick Start, but there are no validation checkpoints and batch operations like sweeps lack verification steps, capping the score at 2 per the feedback-loop guideline.

2 / 3

Progressive Disclosure

Real one-level references are signaled in 'See Also' (sweeps.md, artifacts.md, integrations.md all exist), but large sweeps/artifacts/integrations sections are duplicated inline — content that should be split into the reference files.

2 / 3

Total

9

/

12

Passed

Description

82%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A terse, specific description that names concrete W&B capabilities and is clearly distinguishable, but it lacks an explicit 'Use when…' trigger clause, capping completeness at 2.

Suggestions

Add an explicit trigger clause, e.g. 'Use when logging ML experiments, running hyperparameter sweeps, or managing a model registry with Weights & Biases.'

Include the full product name 'Weights & Biases' alongside 'W&B' so the description matches users who spell it out.

Lead with a verb for each capability (e.g. 'run sweeps', 'manage model registry') rather than bare nouns to strengthen the action framing.

DimensionReasoningScore

Specificity

Enumerates multiple concrete capabilities — "log ML experiments, sweeps, model registry, dashboards" — rather than vague language, matching the 'lists multiple specific concrete actions' anchor.

3 / 3

Completeness

It clearly states what the skill does but omits any 'Use when…' trigger, so per the judging guideline a missing explicit trigger clause caps completeness at 2 rather than 3.

2 / 3

Trigger Term Quality

Natural user-facing terms (W&B, ML experiments, sweeps, model registry, dashboards) give good coverage of what a practitioner would say; not merely technical jargon, so it clears the level-2 bar.

3 / 3

Distinctiveness Conflict Risk

The W&B-specific niche and named features make it clearly distinguishable and unlikely to trigger for the wrong skill; not generic like 'helps with code and documents'.

3 / 3

Total

11

/

12

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (595 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.