CtrlK
BlogDocsLog inGet started
Tessl Logo

saelens

Train sparse autoencoders to interpret model features.

55

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/saelens/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is highly actionable and well-structured, with executable workflows, checklists, and a clean one-level reference layout. Its main weakness is conciseness: conceptual primers and marketing-flavored asides occupy tokens that do not earn their place for an audience that already knows SAE fundamentals.

Suggestions

Trim or move the "Problem: Polysemanticity & Superposition" primer and the Anthropic-research feature list into references/, keeping only what motivates a concrete decision.

Remove marketing language ("1,100+ stars", "groundbreaking research") that adds no actionable signal.

Add an explicit validate→fix→retry feedback loop to the training workflow (e.g., check L0/dead-feature metrics after a checkpoint, adjust L1 or warm-up, resume) to push workflow clarity to 5.

DimensionReasoningScore

Conciseness

The body is mostly efficient with concrete code and tables, but retains explanations Claude already knows (the polysemanticity/superposition primer), marketing padding ("1,100+ stars", "groundbreaking research"), and a padded feature example list (DNA, Hebrew text, nutrition statements) that could be trimmed.

3 / 5

Actionability

Three complete, copy-paste-ready workflows (loading, training, analysis/steering) plus WRONG-vs-RIGHT issue fixes, executable configs, and concrete metrics tables fully cover the common cases.

5 / 5

Workflow Clarity

Each workflow uses numbered Step-by-Step sequences with closing checklists and evaluation metrics as checkpoints, and the Common Issues section supplies error-recovery patterns, but there are no explicit validate→fix→retry feedback loops in the main flows.

4 / 5

Progressive Disclosure

SKILL.md serves as an overview with a clearly signaled, one-level-deep Reference Documentation table pointing to real bundle files (README.md, api.md, tutorials.md), though a fair amount of detail (full class reference, common issues, architecture tables) remains inline rather than split out.

4 / 5

Total

16

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names a specific, distinctive domain with a concrete primary action, but it lacks any "when to use" trigger guidance and misses common synonyms (e.g., SAE, mechanistic interpretability). These gaps cap completeness and trigger-term quality at the midpoint.

Suggestions

Add an explicit trigger clause, e.g. "Use when training or analyzing sparse autoencoders (SAEs) for mechanistic interpretability or feature discovery."

Include common synonyms users would naturally say — "SAE", "mechanistic interpretability", "feature discovery", "superposition" — to improve trigger coverage.

Broaden the action list beyond "Train" to reflect the skill's workflows, e.g. "Train, load, and analyze sparse autoencoders to discover interpretable model features."

DimensionReasoningScore

Specificity

Names the domain ("sparse autoencoders") with one concrete action ("Train") and a goal ("interpret model features"), but does not enumerate the broader capability set (load pre-trained SAEs, analyze, steer) so coverage is not comprehensive.

3 / 5

Completeness

Provides a clear "what" (train SAEs to interpret features) but includes no "Use when…" clause or equivalent explicit trigger guidance, which caps completeness at 3 per the rubric guidelines.

3 / 5

Trigger Term Quality

Contains relevant keywords ("sparse autoencoders", "model features", "interpret") but omits common synonyms and variations a user would naturally say, such as "SAE", "mechanistic interpretability", or "feature discovery".

3 / 5

Distinctiveness Conflict Risk

The sparse-autoencoder niche is mostly distinct with only minor overlap risk against general interpretability tools, but the absence of explicit trigger phrases keeps it just below a clearly-distinct 5.

4 / 5

Total

13

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.