CtrlK
BlogDocsLog inGet started
Tessl Logo

skillopt-sleep

Use when the user wants their Claude agent to self-improve from past usage, asks about a nightly/offline 'sleep' or 'dream' cycle, memory/skill consolidation, or says things like 'make my agent better the more I use it', 'review my past sessions', 'learn my preferences', 'consolidate what you learned', 'run the sleep cycle', or wants to schedule offline self-optimization. Drives the skillopt_sleep engine: harvest past sessions -> mine recurring tasks -> replay offline -> consolidate validated CLAUDE.md and SKILL.md behind a held-out gate.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is skillopt-sleep in alirezarezvani/claude-skills

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable and the workflow is exemplary — a clearly sequenced six-stage cycle with explicit held-out validation, staging, and backups guarding the destructive adopt step. The main slack is conceptual exposition in the intro and a lack of any references/ split for the flag-table and config detail.

Suggestions

Cut the "It synthesizes three ideas" section and the "deployment-time analogue of training" framing down to one or two lines, or move that background to the skill's README.md.

Move the full CLI flag table and the config-keys list into a references/ file (e.g. references/cli.md), keeping the four core commands (status/dry-run/run/adopt) inline in SKILL.md.

Drop the closing meta-note about the original design-doc path not being vendored; keep only the upstream guide URL, and move the vendoring explanation to README.md.

DimensionReasoningScore

Conciseness

The operational core is lean (commands, flag table, config keys), but the conceptual intro ("the deployment-time analogue of training: short-term experience → long-term competence") and the "synthesizes three ideas" section (SkillOpt / Claude Dreams / Agent sleep) explain background Claude does not need to run the skill. This is minor, localized over-explanation — the 4 anchor — rather than the pervasive padding of 3.

4 / 5

Actionability

Commands are copy-paste ready across all operations: "${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" status/dry-run/run/adopt, schedule/unschedule, plus a full 14-row flag table with defaults, config keys, and a deterministic validation command (python -m skillopt_sleep.experiments.run_experiment --persona researcher --assert-improves). Matches the 5 anchor for fully executable coverage of common cases.

5 / 5

Workflow Clarity

The six-stage cycle (Harvest -> Mine -> Replay -> Consolidate -> Stage -> Adopt) is clearly sequenced, and validation around the destructive adopt step is explicit: held-out gate, staging where "Nothing live changes", backups on adopt, and "Always show the user the held-out baseline → candidate score ... Evidence before adoption." This matches the 5 anchor with explicit validation and error-recovery guidance (dry-run).

5 / 5

Progressive Disclosure

Sections are well organized and the skill is effectively self-contained (no references/ bundle exists; external detail is one-level-deep via the upstream guide URL). However, the full CLI flag table and config-key reference are inlined where a references/ file would serve better, and the closing meta-note about the non-vendored design-doc path is README material. Fits the 4 anchor: good structure, minor organization gaps.

4 / 5

Total

18

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it answers both what and when with concrete, third-person capability statements and an unusually rich set of natural trigger phrases. The only weakness is mild overlap risk with general memory/CLAUDE.md-management skills on a few trigger terms.

DimensionReasoningScore

Specificity

The description lists a concrete action pipeline — "harvest past sessions -> mine recurring tasks -> replay offline -> consolidate validated CLAUDE.md and SKILL.md behind a held-out gate" — giving comprehensive, specific coverage of what the engine does, in third-person voice. It matches the 5 anchor rather than 4 because the what is covered completely, with no notable gaps.

5 / 5

Completeness

Both halves are explicit: "Use when the user wants their Claude agent to self-improve... or says things like..." answers when with concrete trigger phrases, and "Drives the skillopt_sleep engine: harvest -> mine -> replay -> consolidate" answers what. This is a direct match for the 5 anchor.

5 / 5

Trigger Term Quality

It includes a wide range of natural user phrasings: "make my agent better the more I use it", "review my past sessions", "learn my preferences", "consolidate what you learned", "run the sleep cycle", plus synonyms like nightly/offline "sleep" or "dream" cycle. Coverage is comprehensive with variants, matching the 5 anchor.

5 / 5

Distinctiveness Conflict Risk

The sleep/dream-cycle framing and skillopt_sleep naming are a clear niche, but phrases like "learn my preferences", "memory/skill consolidation", and "consolidate ... CLAUDE.md" carry minor overlap risk with generic memory-management skills. Fits the 4 anchor (mostly distinct, minor overlap with closely related skills) better than 5.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
alirezarezvani/claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.