CtrlK
BlogDocsLog inGet started
Tessl Logo

skillopt-sleep

Use when the user wants their Claude agent to self-improve from past usage, asks about a nightly/offline 'sleep' or 'dream' cycle, memory/skill consolidation, or says things like 'make my agent better the more I use it', 'review my past sessions', 'learn my preferences', 'consolidate what you learned', 'run the sleep cycle', or wants to schedule background self-optimization. Drives the skillopt_sleep engine: harvest past sessions -> mine recurring tasks -> replay through a selected backend -> consolidate validated CLAUDE.md/SKILL.md behind a held-out gate.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Highly actionable content with copy-paste CLI commands, a clearly sequenced six-stage workflow, and strong validation/backoff safeguards around a destructive adopt step. It is only mildly held back by trimmable conceptual background and a monolithic single-file layout that inlines reference-grade detail (full flag and config tables).

Suggestions

Trim the conceptual framing (the three-idea synthesis paragraph and 'deployment-time analogue of training' line) and drop or compress the 'When to use this skill' section, since it duplicates the frontmatter description's trigger phrases.

Move the full CLI flag table and config-keys detail into a references/ file (e.g. references/cli.md), keeping a short quick-start set of the most common flags inline in SKILL.md.

DimensionReasoningScore

Conciseness

Mostly efficient: the six-stage cycle, CLI commands, flag table, and config keys each earn their tokens. But the three-idea synthesis ("It is the deployment-time analogue of training: short-term experience -> long-term competence") and the 'When to use' section repeating the description's trigger phrases are trimmable. Anchor 4 (efficient, minor instances that could be trimmed) fits; not 5, not 3.

4 / 5

Actionability

Fully executable, copy-paste-ready commands: "${CLAUDE_PLUGIN_ROOT}/scripts/sleep.sh" status/dry-run/run/adopt, scheduling commands, the demo experiment invocations, a complete flag table with defaults, and concrete config keys. Common cases (status check, full run, adopt, schedule, no-API demo) are all specifically covered.

5 / 5

Workflow Clarity

The six stages (Harvest -> Mine -> Replay -> Consolidate -> Stage -> Adopt) are clearly sequenced, and this risky batch operation (overwriting live CLAUDE.md/SKILL.md) has explicit validation checkpoints: the held-out gate ("accept only if it strictly improves"), staging with "Nothing live changes", backup before adopt, and "Evidence before adoption". Feedback on rejection is also specified (rejected runs still produce a report).

5 / 5

Progressive Disclosure

A single-file skill (no references/, scripts/, or assets/ bundle exists) with well-organized sections and no nested references; the one external pointer (GitHub docs URL) is clearly signaled. However, at ~155 lines the full CLI flag table and config-key details are inlined where a reference file could keep the SKILL.md overview lean, so anchor 4 fits better than 5.

4 / 5

Total

18

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use when...' trigger clause, abundant natural trigger phrasings, and a concrete what-description of the engine pipeline. The only weakness is mild overlap risk with generic memory/preference skills on a couple of trigger phrases.

DimensionReasoningScore

Specificity

"Drives the skillopt_sleep engine: harvest past sessions -> mine recurring tasks -> replay through a selected backend -> consolidate validated CLAUDE.md/SKILL.md behind a held-out gate" lists multiple concrete actions with named artifacts (CLAUDE.md, SKILL.md, held-out gate), fully covering the pipeline. Not 4: there are no minor gaps in coverage of what the skill does.

5 / 5

Completeness

Explicitly answers both: what ("Drives the skillopt_sleep engine: harvest... consolidate validated CLAUDE.md/SKILL.md behind a held-out gate") and when ("Use when the user wants their Claude agent to self-improve from past usage...") with concrete trigger phrases. Matches the score-5 anchor example structure.

5 / 5

Trigger Term Quality

Includes comprehensive natural phrasings users would actually say: "'make my agent better the more I use it'", "'review my past sessions'", "'learn my preferences'", "'consolidate what you learned'", "'run the sleep cycle'", plus synonyms (sleep, dream, consolidation, self-optimization). Not 4: synonyms and colloquial variations are already covered.

5 / 5

Distinctiveness Conflict Risk

The agent-self-improvement/sleep-cycle niche is clear with distinct triggers, but "memory/skill consolidation" and "'learn my preferences'" overlap with what a generic memory-management skill would claim. Anchor 4 ("minor overlap risk with closely related skills") fits better than 5 ("minimal conflict risk").

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
microsoft/SkillOpt
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.