CtrlK
BlogDocsLog inGet started
Tessl Logo

weights-and-biases

Track ML experiments with automatic logging, visualize training in real-time, optimize hyperparameters with sweeps, and manage model registry with W&B - collaborative MLOps platform

52

Quality

60%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/weights-and-biases/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with concrete, executable code for every major W&B feature, but it substantially duplicates its own reference bundle (a ~590-line SKILL.md alongside 2,100 lines of references) and lacks validation/feedback loops for batch operations like sweeps. Restructuring to an overview that defers detail to the references and adds sweep monitoring guidance would address the main weaknesses.

Suggestions

Cut the inlined integration examples and sweep-strategy details down to one short pointer each (e.g., "**Integrations**: See references/integrations.md") since dedicated reference files already cover them, and remove the Pricing/Resources/Team Collaboration padding.

Link each reference file at the relevant section (Artifacts → references/artifacts.md, Sweeps → references/sweeps.md) rather than only in a terminal "See Also" block, so navigation is signaled where the reader needs it.

Add validation/feedback steps for sweep runs — how to check sweep status (wandb sweep --show, the UI URL), what to do when trials fail, and when to stop early — to lift workflow clarity for batch operations.

DimensionReasoningScore

Conciseness

The ~590-line body is noticeably verbose: it inlines full integration examples (HuggingFace, Lightning, Keras) and detailed sweep strategies despite dedicated reference files (references/integrations.md, references/sweeps.md), and pads with sections like Pricing, Team Collaboration, and Resources. Not 1: it does not explain concepts Claude already knows (no "what is ML" padding) and is mostly concrete code; not 3: the duplication with the bundle files and the non-instructional sections go beyond minor tightening.

2 / 5

Actionability

Mostly executable, copy-paste-ready examples for the common cases — wandb.init with config, wandb.log variants, sweep config dicts, artifact logging, and framework callbacks. Not 5: several examples reference undefined helpers (train_epoch(), build_model(), get_optimizer(), train_acc) so they are not fully runnable as-is; not 3: the gaps are minor and the code is real rather than pseudocode.

4 / 5

Workflow Clarity

Sequences are present and coherent (credential check → install/login → init → log → finish; define sweep → init → agent), and there is one validation step (checking WANDB_API_KEY). Not 4: the batch sweep operation (wandb.agent(sweep_id, function=train, count=50)) has no verification or feedback loop — no guidance on monitoring sweep progress, checking run status, or handling failed trials — and the rubric caps batch operations without validation at 3.

3 / 5

Progressive Disclosure

The body has good section structure and a "See Also" listing three real, one-level-deep reference files, but substantial content that belongs in those files (integration examples, sweep strategy details) is inlined in SKILL.md, and the references are only linked in a terminal "See Also" section rather than at the relevant sections. Not 4: the inlining is significant, not minor; not 2: the references exist, are real files, contain no nested references, and are clearly signaled.

3 / 5

Total

12

/

20

Passed

Description

71%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, capability-dense description that names four concrete actions and a distinctive platform anchor, but it omits any "Use when..." trigger guidance, which caps its completeness and limits synonym coverage for natural user phrasings. Adding an explicit when-to-use clause and common variations like "wandb" or "experiment tracking" would lift it into the top band.

Suggestions

Add a "Use when..." clause, e.g., "Use when the user mentions W&B, wandb, experiment tracking, or wants to log, compare, or sweep ML training runs."

Include natural synonyms and variations users actually say — "wandb", "experiment tracking", "training curves", "metrics logging" — to improve trigger term coverage.

Tighten the generic tail phrase "collaborative MLOps platform" or fold it into the trigger clause so it does not broaden overlap with competing tracking tools.

DimensionReasoningScore

Specificity

The description lists four concrete, non-overlapping actions — "Track ML experiments with automatic logging", "visualize training in real-time", "optimize hyperparameters with sweeps", "manage model registry" — which comprehensively cover the skill's scope. Not 4: there are no minor gaps in coverage; each capability named in the body maps to a phrase here.

5 / 5

Completeness

The "what" is clear and concrete (tracking, visualization, sweeps, registry), but there is no "Use when..." clause or equivalent explicit trigger guidance anywhere in the description — per the judging guidelines this caps completeness at 3. Not 4: the "when" is entirely absent rather than merely implicit or under-specified.

3 / 5

Trigger Term Quality

Good keyword coverage with natural terms users would say ("ML experiments", "hyperparameters", "sweeps", "model registry", "W&B"), but it misses common synonyms and variations such as "wandb", "experiment tracking", "metrics", or "training curves". Not 5: no synonyms or extension-style variations are present; not 3: multiple relevant natural keywords are covered, not just a couple of generic ones.

4 / 5

Distinctiveness Conflict Risk

"with W&B" anchors a clear niche and is a distinct trigger, but the generic lead phrasing ("Track ML experiments", "collaborative MLOps platform") overlaps with other experiment-tracking skills (e.g., MLflow, TensorBoard) until the W&B qualifier is reached. Not 5: minor overlap risk with closely related MLOps tracking skills remains; not 3: the platform name makes it mostly distinct.

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (611 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.