CtrlK
BlogDocsLog inGet started
Tessl Logo

tensorboard

Visualize training metrics, debug models with histograms, compare experiments, visualize model graphs, and profile performance with TensorBoard - Google's ML visualization toolkit

52

Quality

60%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/tensorboard/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, executable reference with a sensible section order and real companion files, but it underperforms on token efficiency and disclosure structure: it duplicates its own Quick Start patterns, inlines ~600 lines of material that the references already cover in more depth, and hides those references in a footer. Adding point-of-use reference links and a verification step (confirm metrics render at localhost:6006, flush/close the writer) would lift the two weakest dimensions.

Suggestions

Cut the duplicated SummaryWriter/scalar-logging sections and the full generic training loops; keep one minimal Quick Start per framework and delegate depth to the references files.

Replace the footer 'See Also' with inline pointers at each section (e.g. in Performance Profiling: "See [references/profiling.md](references/profiling.md) for memory profiling and bottleneck detection") so references are surfaced where needed.

Add an explicit verify step after launching (open http://localhost:6006 and confirm the run appears; if not, check writer.flush()/close() and the --logdir path) to give the workflow a validation checkpoint.

DimensionReasoningScore

Conciseness

At ~620 lines the body is noticeably verbose: 'Core Concepts 1. SummaryWriter' and 'Logging Scalars' restate the Quick Start pattern with repeated imports and writer setup, and the ~75-line PyTorch training loop plus the Keras model definition are generic training code Claude already knows, matching the 'several unnecessary or padded sections' anchor better than the mostly-efficient one.

2 / 5

Actionability

Nearly all guidance is concrete, copy-paste-ready code and commands (SummaryWriter calls, TensorBoard callback flags, `tensorboard --logdir=runs`), with only minor gaps: undefined `file_writer` in the TF histogram/image blocks, `make_grid` not imported in the integration example, placeholder `train_epoch()`/`validate()` calls, and `torch.stack([img1, img2, ...])` fragments — consistent with the 'mostly executable with minor gaps' anchor, not a 5.

4 / 5

Workflow Clarity

The install → create writer → log → launch → view sequence is present but only implicit (per-section "Launch: tensorboard --logdir=runs" hints), with no explicit checkpoints or verification/troubleshooting for the common failure modes (data not appearing because the writer wasn't flushed/closed or the logdir is wrong). Logging is not destructive, so the batch/destructive cap does not apply, and anchor 3 ('sequence present but checkpoints missing or implicit') fits better than 4.

3 / 5

Progressive Disclosure

Three real, one-level-deep reference files exist (references/visualization.md, profiling.md, integrations.md — verified on disk), but they are only signaled in a 'See Also' list at the very end rather than at the point of need, and the SKILL.md inlines substantial material (e.g. its Performance Profiling section overlaps references/profiling.md) instead of acting as an overview — matching anchor 3 ('references present but not clearly signaled; content that should be separate is inline').

3 / 5

Total

12

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A concrete, well-scoped description that names the tool and five specific capabilities with strong natural trigger terms. Its main defect is the absent 'when to use' clause, which caps completeness, plus minor keyword gaps (framework names, 'loss curves', 'experiment tracking').

Suggestions

Append an explicit trigger clause, e.g. "Use when logging or visualizing ML training runs, comparing experiments, or when the user mentions TensorBoard, SummaryWriter, or loss/metric curves."

Add natural synonyms users actually say — "PyTorch/TensorFlow", "experiment tracking", "training curves", "runs/logdir" — to strengthen trigger coverage.

Mention the omitted headline capabilities (embedding projector, hyperparameter comparison) either in the description or trim them from the body's 'When to Use' list so both stay consistent.

DimensionReasoningScore

Specificity

The description lists five concrete actions ("Visualize training metrics", "debug models with histograms", "compare experiments", "visualize model graphs", "profile performance") anchored to a named tool, matching the 'several specific actions with minor gaps' anchor rather than a 5 — the skill's own body advertises capabilities it omits (embedding projection, hyperparameter tracking, image/text logging).

4 / 5

Completeness

The 'what' is clear and multi-part, but there is no 'Use when...' clause or equivalent trigger guidance, which per the judging guidelines caps completeness at 3 ("Has a clear 'what' but 'when' is missing").

3 / 5

Trigger Term Quality

Natural terms users would actually say are present ("training metrics", "compare experiments", "debug models", "histograms", "profile performance", "TensorBoard"), but common variations are missing — no "PyTorch"/"TensorFlow", "loss curves", "experiment tracking", or "runs/logdir", so it falls between the good-coverage (4) and comprehensive (5) anchors.

4 / 5

Distinctiveness Conflict Risk

Naming TensorBoard explicitly with its distinctive feature set carves out a clear niche with minimal overlap risk; no competing logging/visualization skill would be wrongly triggered by these phrases.

5 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (639 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.