CtrlK
BlogDocsLog inGet started
Tessl Logo

sparse-autoencoder-training

Provides guidance for training and analyzing Sparse Autoencoders (SAEs) using SAELens to decompose neural network activations into interpretable features. Use when discovering interpretable features, analyzing superposition, or studying monosemantic representations in language models.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured skill body with executable workflows, checklists, metric-based validation, and properly signaled one-level references. The main drag on token efficiency is the conceptual background and promotional fluff that Claude does not need.

Suggestions

Cut the 'The Problem: Polysemanticity & Superposition' and 'What SAEs Learn / Key Validation' preamble plus the '1,100+ stars' note; assume Claude knows SAE basics and keep only SAELens-specific guidance.

Move the conceptual Anthropic-research context (70% interpretable, feature examples) into references/ so the body stays a lean operational guide.

DimensionReasoningScore

Conciseness

Mostly efficient with strong code blocks, but the conceptual preamble ('The Problem: Polysemanticity & Superposition', 'What SAEs Learn', the 70%-interpretable Anthropic stat) and fluff like '1,100+ stars' explain background Claude already knows and could be trimmed.

2 / 3

Actionability

Workflows are built from complete, executable, copy-paste-ready code (loading/encoding, full LanguageModelSAERunnerConfig training, steering hooks, logit attribution) with concrete parameter values, matching the fully-executable anchor.

3 / 3

Workflow Clarity

Each workflow is numbered step-by-step with a closing checklist, and the training workflow includes verification via metric targets (L0 50-200, CE 80-95%, dead features <5%) plus a wrong/right troubleshooting section — giving explicit validation checkpoints for the batch training operation.

3 / 3

Progressive Disclosure

The body is an overview with three workflows inline and a clearly signaled, one-level-deep reference table pointing to real files (references/README.md, api.md, tutorials.md) for detail, matching the well-organized one-level reference anchor.

3 / 3

Total

11

/

12

Passed

Description

85%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with concrete capabilities, an explicit use-when trigger, and a distinct mech-interpretability niche. Its main weakness is trigger-term naturalness: the phrasing is jargon-heavy and could miss how a user actually phrases the need.

Suggestions

Soften jargon in the trigger clause by adding natural phrasings users actually say, e.g. 'Use when training or finding interpretable features in a model with sparse autoencoders (SAEs), analyzing superposition, or exploring monosemantic features.'

Add a common-variation keyword like 'feature discovery' or 'find what a neuron represents' alongside 'discovering interpretable features' to broaden natural trigger coverage.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'training and analyzing Sparse Autoencoders', 'decompose neural network activations into interpretable features' — matching the multi-action anchor rather than a single vague verb.

3 / 3

Completeness

Explicitly states what it does ('training and analyzing... to decompose... activations') and when to use it via a 'Use when discovering interpretable features, analyzing superposition, or studying monosemantic representations' trigger clause.

3 / 3

Trigger Term Quality

Includes relevant niche keywords ('interpretable features', 'superposition', 'SAEs') but leans on jargon like 'monosemantic representations' and omits common variations a user might actually say ('find features', 'train an SAE'), so it stops short of full natural-term coverage.

2 / 3

Distinctiveness Conflict Risk

The SAELens / Sparse Autoencoder scope with mech-interp-specific triggers carves a clear niche unlikely to fire for unrelated skills.

3 / 3

Total

11

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
Orchestra-Research/AI-Research-SKILLs
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.