CtrlK
BlogDocsLog inGet started
Tessl Logo

nnsight-remote-interpretability

Provides guidance for interpreting and manipulating neural network internals using nnsight with optional NDIF remote execution. Use when needing to run interpretability experiments on massive models (70B+) without local GPU resources, or when working with any PyTorch architecture.

57

Quality

67%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/nnsight/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with five concrete, mostly copy-paste-ready workflows and clearly signaled reference files that all exist. Its weaknesses are padding in the marketing/comparison/external-links sections and substantial inline detail that duplicates the references/ bundle.

Suggestions

Trim 'Key Value Proposition', 'Architecture Support', and the 'Comparison with Other Tools'/'External Resources' sections down to a line or two each, moving detail into references/.

Make the 'Systematic Patching Sweep' self-contained by defining clean_cache, seq_len, and compute_metric or explicitly labeling them as user-supplied.

Move the bulk of the workflow walkthroughs and the API table into tutorials.md and api.md, keeping only one quick-start example per pattern in SKILL.md.

DimensionReasoningScore

Conciseness

The workflow code is tight, but 'Key Value Proposition', 'Architecture Support', the 'Comparison with Other Tools' table, and the long link lists in 'External Resources' add material Claude could infer or that duplicates the reference files, so the body could be noticeably tightened.

3 / 5

Actionability

Nearly all code blocks are complete and executable (loading, tracing, saving, patching, remote execution), but the 'Systematic Patching Sweep' depends on undefined clean_cache, seq_len, and compute_metric, leaving a minor gap.

4 / 5

Workflow Clarity

Each workflow is sequenced with numbered in-code steps and Workflow 1 has an explicit checklist, but only that one workflow has a checklist and no workflow includes an explicit verify-output step, leaving minor validation gaps.

4 / 5

Progressive Disclosure

The references table signals real one-level-deep files (references/README.md, api.md, tutorials.md) with content descriptions, but the body inlines substantial overlapping detail (an API reference table despite api.md, five full workflows despite tutorials.md), making it long for an overview.

4 / 5

Total

15

/

20

Passed

Description

67%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description answers both what and when with an explicit, concrete 'Use when...' clause, in proper third-person voice. It is held back by a generic action list ('Provides guidance', 'run interpretability experiments') and a very broad 'any PyTorch architecture' trigger that raises conflict risk.

Suggestions

Replace generic phrasing with concrete capability verbs, e.g. 'Traces, saves, and patches activations in any PyTorch model; runs the same code remotely on 70B+ models via NDIF.'

Add natural trigger synonyms users would actually say, such as 'mechanistic interpretability', 'activation patching', or 'model internals'.

Narrow the 'when working with any PyTorch architecture' trigger to internals-access contexts (e.g. 'when you need to inspect or intervene on activations inside a PyTorch model') to reduce overlap with generic PyTorch skills.

DimensionReasoningScore

Specificity

Names the domain ('interpreting and manipulating neural network internals using nnsight') and a couple of actions, but phrases like 'Provides guidance' and 'run interpretability experiments' stay generic without concrete operations such as activation patching, gradient analysis, or cross-prompt sharing.

3 / 5

Completeness

The what is explicit (interpreting and manipulating neural network internals via nnsight with NDIF remote execution) and the when is stated with concrete trigger phrases: 'Use when needing to run interpretability experiments on massive models (70B+) without local GPU resources, or when working with any PyTorch architecture.'

5 / 5

Trigger Term Quality

Relevant keywords are present ('nnsight', 'NDIF', 'remote execution', 'interpretability experiments', '70B+', 'PyTorch'), but natural variations users would say — 'mechanistic interpretability', 'activation patching', 'model internals', 'probing' — are missing.

3 / 5

Distinctiveness Conflict Risk

The nnsight/NDIF niche is distinct, but the trigger 'or when working with any PyTorch architecture' is extremely broad and would fire on nearly any PyTorch task, creating real overlap risk with generic PyTorch/training skills.

3 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.