CtrlK
BlogDocsLog inGet started
Tessl Logo

weights-and-biases

Track ML experiments with automatic logging, visualize training in real-time, optimize hyperparameters with sweeps, and manage model registry with W&B - collaborative MLOps platform

61

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./bundled/skills/weights-and-biases/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with executable examples and real, well-signaled reference files, but it is verbose and duplicates reference content inline while lacking validation checkpoints. Slimming the inlined sweep/artifact/integration sections into the existing references and adding light verification steps would raise the score.

Suggestions

Move the full sweep config, training function, and strategy examples into references/sweeps.md, keeping only a minimal inline snippet and a pointer, to fix the inline/reference duplication.

Trim the marketing line ("200,000+ users / 10.5k+ stars / 100+ integrations") and the time-sensitive Pricing section, or relocate pricing to a reference so the body stays lean and non-stale.

Add explicit validation checkpoints where operations can fail (e.g. confirm wandb login/API key before runs, verify artifact download succeeded before loading, check sweep agent results) to give workflows feedback loops.

DimensionReasoningScore

Conciseness

The ~400-line body is code-heavy and mostly actionable, but it inlines full sweep configs, artifact examples, and HF/Lightning/Keras sections that duplicate the reference files, plus a marketing line and a time-sensitive Pricing section that pad the token budget — tightening and de-duplication would reach the lean score-3 anchor.

2 / 3

Actionability

Code throughout is executable and copy-paste ready (wandb.init, wandb.log, sweep config + agent, Artifacts, and framework integrations), matching the fully-executable score-3 anchor.

3 / 3

Workflow Clarity

Content is organized as a topic catalog with quick-start and sweep sequences, but lacks explicit validation checkpoints or feedback loops — fitting the score-2 "sequence present but checkpoints missing" anchor rather than score 3.

2 / 3

Progressive Disclosure

The three reference files exist and are clearly signaled one level deep in "See Also", but the SKILL.md heavily duplicates sweeps/artifacts/integrations content inline that already lives in those references, fitting the score-2 "content that should be separate is inline" anchor.

2 / 3

Total

9

/

12

Passed

Description

82%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and trigger-rich with a clear W&B niche, but it omits an explicit "Use when…" trigger clause, which caps completeness at 2. Adding when-to-use guidance would lift it to a top-band description.

Suggestions

Append an explicit trigger clause, e.g. "Use when tracking ML experiments, visualizing training, running hyperparameter sweeps, or managing a model registry with W&B."

Spell out "Weights & Biases (wandb)" once so the full natural name and CLI token both appear, improving trigger recall.

Keep the imperative third-person voice (already correct) and drop redundant tokens like "- collaborative MLOps platform" if a trigger clause is added to stay concise.

DimensionReasoningScore

Specificity

"Track ML experiments with automatic logging, visualize training in real-time, optimize hyperparameters with sweeps, and manage model registry with W&B" lists multiple concrete actions, matching the score-3 anchor; voice is imperative/third person so no penalty applies.

3 / 3

Completeness

The "what" is strong but there is no "Use when…" clause or equivalent explicit trigger guidance, so per the judging guidelines completeness is capped at 2 rather than reaching the score-3 both-what-and-when anchor.

2 / 3

Trigger Term Quality

Natural practitioner terms are well covered — "ML experiments", "training", "hyperparameters", "sweeps", "model registry", "W&B", "MLOps" — matching the good-coverage anchor rather than the partial-coverage score-2 anchor.

3 / 3

Distinctiveness Conflict Risk

The W&B/MLOps niche is clearly scoped with distinct triggers ("W&B", "model registry", "sweeps", "MLOps") and is unlikely to fire for an unrelated skill.

3 / 3

Total

11

/

12

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (591 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
foryourhealth111-pixel/Vibe-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.