CtrlK
BlogDocsLog inGet started
Tessl Logo

sparse-autoencoder-training

Provides guidance for training and analyzing Sparse Autoencoders (SAEs) using SAELens to decompose neural network activations into interpretable features. Use when discovering interpretable features, analyzing superposition, or studying monosemantic representations in language models.

63

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/saelens/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-built, highly actionable skill document: complete executable code for all three core workflows, useful hyperparameter and metrics tables, checklists, and a symptom-based troubleshooting section. Its main costs are token spend on background Claude already knows and duplication between the inline workflows and references/tutorials.md, plus one broken bundle link.

Suggestions

Cut or compress the 'The Problem: Polysemanticity & Superposition' and 'Key Validation (Anthropic Research)' sections to one or two lines each — Claude knows this background; keep only the SAELens-specific facts (loss form, metric targets).

Deduplicate the three full inline workflows against references/tutorials.md — keep condensed quick-start versions in SKILL.md and point to tutorials.md for the complete step-by-step scripts.

Fix the bundle navigation: references/README.md lists papers.md, which does not exist — either add it or remove the link.

DimensionReasoningScore

Conciseness

Largely code- and table-driven, but it opens with ~15 lines of background Claude already knows ('polysemantic', 'models use superposition to represent more features than they have neurons', Anthropic research trivia like '1,100+ stars' and the 70%-interpretable statistic) — more than the 'minor instances' of the 4 anchor, yet not the heavily padded prose of the 2 anchor.

3 / 5

Actionability

Three mostly complete, copy-paste-ready workflows (loading/encoding, full training config with all hyperparameters, steering and attribution) plus symptom→fix snippets place this above 'minor gaps' at 3, but fragments like `trainer.metrics['l0']` and the partial LanguageModelSAERunnerConfig snippets in Common Issues are not independently executable, keeping it below the 5 anchor's 'fully executable, covers common cases' bar.

4 / 5

Workflow Clarity

Each workflow has numbered steps, a follow-up checklist, and evaluation metric targets (L0 50–200, CE loss 80–95%, dead features <5%) with a Common Issues troubleshooting section — a clear sequence with most checkpoints present. The 5 anchor requires validation wired into the step sequence as explicit validate→fix→retry loops; here validation lives in declarative checklists and a separate troubleshooting section rather than inline steps.

4 / 5

Progressive Disclosure

Good structure: a dedicated 'Reference Documentation' section with a table clearly signaling the three one-level-deep, existing reference files (README.md, api.md, tutorials.md), with most overview-level content in SKILL.md. It falls short of the 5 anchor because the 375-line body duplicates full tutorial workflows that references/tutorials.md also contains, and references/README.md links a papers.md that does not exist in the bundle.

4 / 5

Total

15

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete capabilities with the library named, and includes an explicit, well-phrased 'Use when...' trigger clause covering the three main use cases. Its weaknesses are mild — slightly generic action framing and minor overlap risk with adjacent interpretability tooling.

Suggestions

Make actions more granular (e.g., 'Load pre-trained SAEs, train custom SAEs, and steer model behavior with discovered features') instead of 'Provides guidance for training and analyzing'.

Add 'mechanistic interpretability' and 'feature steering' as trigger terms, and consider scoping the 'when' clause so it doesn't catch plain TransformerLens activation-analysis requests.

DimensionReasoningScore

Specificity

Names the domain (SAELens) and several concrete actions — 'training and analyzing Sparse Autoencoders', 'decompose neural network activations into interpretable features' — but 'provides guidance for' is somewhat abstract, falling short of the fully granular action list of the 5 anchor.

4 / 5

Completeness

Clearly answers 'what' ('training and analyzing Sparse Autoencoders (SAEs) using SAELens to decompose neural network activations into interpretable features') and an explicit 'Use when discovering interpretable features, analyzing superposition, or studying monosemantic representations' trigger clause with concrete trigger phrases — matching the 5 anchor; a 4 would require a less explicit 'when'.

5 / 5

Trigger Term Quality

Good natural keyword coverage: 'Sparse Autoencoders (SAEs)', 'SAELens', 'interpretable features', 'superposition', 'monosemantic representations' — phrases users in this domain actually say. A few natural variations are missing (e.g., 'mechanistic interpretability', 'feature steering', 'dictionary learning'), so not the comprehensive synonym coverage of the 5 anchor.

4 / 5

Distinctiveness Conflict Risk

The SAELens/SAE niche is highly distinct, but trigger phrases like 'discovering interpretable features' and 'analyzing superposition' could also plausibly match general interpretability skills (TransformerLens, pyvene), giving minor overlap risk — the 5 anchor requires minimal conflict risk.

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.