CtrlK
BlogDocsLog inGet started
Tessl Logo

continuous-learning-v2

フックを介してセッションを観察し、信頼度スコアリング付きのアトミックなインスティンクトを作成し、スキル/コマンド/エージェントに進化させるインスティンクトベースの学習システム。

49

Quality

54%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./docs/ja-JP/skills/continuous-learning-v2/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-organized with a clear quickstart, concrete hook configuration, and useful tables, but it is a monolithic single file that inlines reference material, contains a corrupted json code block (prose where the plugin-install hook config should be), and offers no validation step to confirm the observation pipeline is actually working. Fixing the broken code block, adding a post-setup verification step, and moving the config reference and confidence tables into reference files would raise most dimensions.

Suggestions

Fix the broken ```json block in the 'プラグインとしてインストールした場合' section — it currently contains Japanese prose where the actual (or 'no configuration needed') instruction/hook config should be, making that install path unexecutable.

Add a validation checkpoint to the quickstart, e.g. 'run any tool call, then confirm ~/.claude/homunculus/observations.jsonl has grown / run /instinct-status to verify observation is live', since editing settings.json hooks is a risky silent-failure configuration change.

Move the full config.json reference, confidence-scoring tables, and the ASCII pipeline diagram into a references/ file (e.g. ARCHITECTURE.md) and trim the v1/v2 marketing table and quote section to reclaim token budget.

DimensionReasoningScore

Conciseness

The core instructions are domain-specific and lean, but the v1/v2 comparison table, the promotional quote section ('スキルは...約50-80%の確率で発火します' / 'フックは100%の確率で...発火します'), the closing tagline, and inline version-sensitive text ('Claude Code v2.1+') are padding that could be trimmed. Matches anchor 3; not 4 because several sections over-explain, not 2 because the operational content is genuinely efficient.

3 / 5

Actionability

Concrete executable guidance exists (the manual-install hooks JSON block, 'mkdir -p ~/.claude/homunculus/...' commands, the /instinct-* command table), but the 'プラグインとしてインストールした場合' branch places Japanese prose inside a ```json fence instead of actual configuration, and referenced artifacts (hooks/observe.sh, hooks/hooks.json, the Python CLI) are not present in the skill bundle. Matches anchor 3; not 4 because a broken executable block and missing files prevent copy-paste success, not 2 because most guidance is specific and executable.

3 / 5

Workflow Clarity

The quickstart is a clear numbered 1-2-3 sequence (enable hooks, initialize directories, use commands), but there is no validation checkpoint anywhere — nothing to confirm hooks actually fire or that observations.jsonl is being written after the risky settings.json edit. Matches anchor 3; not 4 because checkpoints are absent rather than merely minor, not 2 because the sequence itself is well-ordered and complete.

3 / 5

Progressive Disclosure

Section headers are consistent and navigation within the single file is easy, but this is a ~260-line monolith that inlines the full config.json reference, confidence-scoring tables, and a large ASCII pipeline diagram — content that belongs in separate reference files — and the body points to hooks/ and CLI paths that do not exist in the bundle. Matches anchor 3; not 4 because substantial reference material is inline with no external split, not 2 because the internal structure is real and clearly signaled.

3 / 5

Total

12

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states a clear, multi-action 'what' with reasonably concrete verbs (observe, create, evolve) and a distinct niche, but it entirely lacks any 'when to use it' trigger guidance and leans on internal jargon ('instinct', 'confidence scoring') rather than natural user phrasing. Adding an explicit 'Use when...' clause with user-side trigger terms would lift both completeness and trigger quality.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user wants Claude to learn from their corrections and repeated workflows across sessions, or mentions learning, habits, patterns, or preferences.'

Replace or gloss internal jargon ('インスティンクト') with natural user-side terms such as 'learned behaviors', 'patterns', and 'preferences' so the description matches what users actually say.

Briefly state what the system does end-to-end for the user (e.g. 'applies learned behaviors automatically as confidence grows') so the 'what' is comprehensive, not just architectural.

DimensionReasoningScore

Specificity

Lists several concrete actions — "フックを介してセッションを観察し" (observes sessions via hooks), "信頼度スコアリング付きのアトミックなインスティンクトを作成し" (creates atomic instincts with confidence scoring), "スキル/コマンド/エージェントに進化させる" (evolves them into skills/commands/agents) — which matches the 'several specific actions, minor gaps' anchor. Not 5: coverage is not comprehensive (how observation or evolution works is unstated); not 3: it names more than 1-2 actions.

4 / 5

Completeness

The 'what' is clearly stated (observe sessions via hooks, create confidence-scored instincts, evolve them into skills/commands/agents) but there is no 'when to use it' clause or equivalent trigger guidance anywhere, which caps completeness at 3 per the judging guidelines. Not 4: an explicit or even weakly implied 'when' is entirely absent; not 2: the 'what' half is concrete, not vague.

3 / 5

Trigger Term Quality

Contains some relevant keywords (learning, hooks, sessions, skills) but they are system-internal jargon ("インスティンクト", "信頼度スコアリング") rather than natural phrases a user would say, such as 'learn from my corrections' or 'remember my preferences'. Matches anchor 3; not 4 because common variations and natural user terms are missing, not 2 because the keywords are domain-relevant rather than purely generic.

3 / 5

Distinctiveness Conflict Risk

The hook-driven instinct pipeline with confidence scoring is a clear niche unlikely to fire for unrelated skills, matching the 'mostly distinct, minor overlap risk' anchor. Not 5: 'instinct-based learning system' still broadly overlaps with general memory/learning/preference skills; not 3: the pipeline description is more specific than 'somewhat specific'.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.