Monitors task execution for skill improvement opportunities. Use during ANY multi-step task, agentic workflow, or work session where the agent uses tools and produces deliverables. Captures patterns, user corrections, workflow insights, and methodology worth preserving as reusable skills. Also triggers in post-task feedback discussions and when the user mentions skill observations, improvements, the observation log, skill taxonomy, or asks the agent to watch for skill opportunities. Also known as "One Skill to Rule Them All" — trigger on this phrase too. IMPORTANT: invoke this skill before the FIRST tool call of any session and before writing or proposing a plan — any turn that will involve a tool call counts, however simple the opener looks. This sentence is the session-start trigger and the only activation layer that survives an unreachable config file; pair it with a CLAUDE.md instruction or a harness session-start hook (references/environments.md) — description matching alone is not enforceable.
66
83%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Load this for setup questions, compaction/resume behaviour, or when running in an environment without filesystem access.
Four activation tiers exist, and only the strongest is enforced. Pick knowingly; do not assume any of the middle ones is a guarantee.
ls the
workspace path; if it fails, call the folder-picker tool; never state
whether the folder is connected without that probe" — because a config
file inside the workspace folder cannot fire in the very case it
targets: when no folder is connected, there is no config file. Verify
the channel once by asking a folder-less session to quote its
preferences block; a line that does not reach the agent is not a tier.request_cowork_directory), the folder can still be
attached mid-session on request, and the Session Start Protocol should
be run the moment it is — only the part of the session before the
mount is lost. Ask early, but do not treat a declined or skipped
folder selection as irreversible. If allow_cowork_file_delete or a similar
local-filesystem permission tool is present in the tool surface, the
session runs on the user's machine and the config is reachable; treat
"this session runs in the cloud, so the config was not loaded" as
unverified until the tool surface has been inventoried — and treat the
converse the same way: an environment flag saying a folder is selected
is a hint, and only a probe of the path in the same turn is evidence.SessionStart hook returning
hookSpecificOutput.additionalContext. Even a hook can only inject a
prompt; choosing to invoke a skill remains a model decision. Cowork has
no hook mechanism; there, tiers 1 to 3 are all there is.Word the trigger mechanically, not judgementally. "Task-oriented session" asks the agent to classify the session at its first turn, before it knows how the session will develop, and short factual-looking requests get classified out. The block below keys on tool use instead: any turn that will involve a tool call counts. And activation must precede planning, not merely execution — where a workflow proposes a plan for approval first, a plan written without the relevant skills carries uninformed decisions past the review gate, and approval locks them in. Load skills before exploring, researching or drafting the plan: both, skill first.
Select on the decision, not on the artefact. Selection and loading are separate steps, and the selection step fails in its own way: it runs against what the user handed over — files to read, a link, a screenshot — so every skill whose description names a subject is missed whenever the request's surface form is an activity no skill claims. Before answering, name the DECISION the user is making and match the installed descriptions against that ("the deliverable is advice about X" → load the skill whose description names X), even though the request arrived as a document to review. The high-risk shape is an attachment plus an open question — "here are two files, what do you think?" — which presents as a reading task and hides its own domain; observed, it ran for four substantive turns of domain advice with the domain skill unloaded, while the session-start skills had loaded correctly. That is what makes this distinct from an activation-config failure: the mechanism worked and still selected the wrong set. The failure is silent by construction, because a skill that never loads cannot announce that it was relevant, so the check belongs in the checklist rather than in a disposition.
Before the first tool call of any session — and before writing or
proposing a plan, not merely before executing one — invoke the
task-observer skill AND execute its Session Start Protocol (storage
check, frontmatter scan, review trigger). Loading the skill and running
the protocol are separate steps; a session that loads the file and stops
has activated nothing. Any turn that will involve a tool call counts; do
not classify the session as "too simple" from its opening message.
Select skills on the DECISION the request is about, not on the artefact it
arrived as. Name what the user is deciding, then match the installed skill
descriptions against that — a request handed over as a file to review
still needs the skill whose description names its subject.
After completing each task, check the observation records written this
session and report a one-line summary (ids and titles, or "none logged
and why"). This is the activation backstop: it forces a look at the log,
so a session that silently skipped the protocol is discovered at the
first task boundary instead of never.
Loading a skill is not complete until you have queried the observation
log for OPEN observations naming it and read their bodies:
find "[ABSOLUTE PATH]/skill-observations/observation-log" -maxdepth 1 \
-name '*.md' -exec grep -l "skill:.*<skill-name>" {} +
(Use find, not a bare *.md glob. Under zsh an unmatched glob is an error,
so on an empty log the command never runs and the enclosing block aborts —
and 2>/dev/null does not help, because the redirection belongs to a
command that never starts. This is the first thing that fires on a fresh
install, in every session, for as long as the log is empty.)
Apply their insights to the current work — meaning: let them change what
you do in THIS task. Editing the skill file, or writing the rule into any
other file a later session reads, is acting on the observation and waits
for the review. See "Log, don't act" in SKILL.md: the test is whether the
action leaves a durable change outside the observation log. Run this at every skill load, however many skills load
in one session. The session-start scan does not cover it: that is a
frontmatter sweep over every observation at session start, this is a
body-level lookup for one skill at the moment its rules are applied.
The task-observer workspace for this project is:
[ABSOLUTE PATH]
Every path the skill uses derives from that root and nothing else:
[ABSOLUTE PATH]/skill-observations/observation-log/ (the log)
[ABSOLUTE PATH]/skill-observations/cross-cutting-principles.md
[ABSOLUTE PATH]/skill-updates/ (staging root)
[ABSOLUTE PATH]/skill-updates/PENDING.md (staging manifest)
Never resolve any of them from the current working
directory — a cwd inside an ephemeral checkout (a git worktree, a temporary
clone) is torn down and takes the log with it. Never place the workspace
inside a skills-discovery directory or any path linked into one. If this
environment mints a separate project identity per checkout, or more than
one agent works this project, the pinned path above is the single shared
location; do not derive one per session, tool or project.SKILL.md's [workspace folder] definition states the rule; this is where
to pin it, how a stable anchor still goes plural, and which plural case is
legitimate.
Fill in the path when installing. Pinning turns anchoring into a one-time
decision instead of one the agent re-litigates every session with a fresh
chance to get it wrong. Pin the ROOT, and list the derived paths: a pin
that names only the observation-log directory — the place split-brain was
first noticed — leaves the staging root, the manifest and the principles
file to be re-derived by every session, and parallel sessions resolve
them plausibly and differently (observed: two sessions staged under
skill-observations/skill-updates/, beside the log, while a third staged
under skill-updates/ at the root — three writers, two staging roots,
one manifest that saw half the work). Scope the workspace to what is being observed:
skills installed globally are observed from every project, so their log
must be one absolute path shared across projects and tools — a per-project
or per-tool workspace scatters observations about global skills across
every project the user touches, and a review run in one never sees the
others. Where the environment provides a managed persistence directory
under the project identity (a memory directory it loads every session),
that directory and the identity root are both "stable"; anchor inside the
managed directory when one exists, otherwise on the identity root — never
both.
Scope matches installation scope. Skills installed at user or global
scope are observed from every project, so their log is one user-scope
path shared by every project, tool and agent; a per-project anchor is
right only for skills that exist in one project. A stable anchor can
still be plural — in Claude Code the identity is derived from the
directory the session starts in, so per-subfolder habits shard the log
into several stable anchors that each look like the only one. Decide the
scope once, at install, and pin it. One plural case is legitimate and
is not a shard: logs deliberately kept apart because their observation
bodies carry task context that does not travel, over skills that are
nonetheless installed globally and therefore shared. Read the multi-log
block in weekly-review.md before consolidating anything — it defines
that case and the aggregate review that serves it. A reader who stops at
this paragraph consolidates a set that was meant to stay separate.
The fork rule as usually stated describes the harmless case. "A second empty log beside a populated one is a silent fork" — but an empty log announces itself: the first scan returns nothing, and nobody trusts a backlog of zero. The damaging fork is the one that has accumulated its own entries. From inside it looks exactly like a working log: the scan returns files, the review finds work, every instrument reports health. Nothing in it can reveal that it is a shard.
And pinning does not settle it. Observed: a workspace pinned in a wired session-start hook, correctly configured, running beside a second populated workspace for three days and producing colliding ids. The pin governs the sessions that read it; it does nothing about a session that resolved its anchor before the pin, or by another route.
So the detection has to be positive, and it is nearly free: at the end of
the session-start scan the agent is already holding the skill: values of
every entry. A log whose targets are overwhelmingly user-scope skills,
while the log itself sits under a per-project path, has already
declared the mismatch — two ls and a comparison. Run it when the anchor
is per-project and the targets are not; report the candidate shard rather
than consolidating on your own judgement, because one plural case is
legitimate (see the multi-log block in weekly-review.md).
Before creating a log, search for one. Check the plausible anchor
candidates — the pinned path, the project identity root, the
environment-managed persistence directory, the shared folder, the other
agent's equivalent — for an existing skill-observations/ workspace. If
one exists, adopt it, or consolidate deliberately with the user. A fresh
empty log beside a populated one is a silent fork: both grow
independently, ids collide, and each session sees only half the history.
When consolidating, leave a pointer file at the abandoned location so
sessions anchored there get redirected instead of re-creating the fork.
Config detection (once per session): with filesystem access, check the workspace root's CLAUDE.md (or equivalent) for a task-observer activation instruction — suggest adding it if absent. Suggest; do not create. If no config file exists at all, offer to add one and let the user say yes: creating a tracked project file changes every future session's behaviour, and the governance-protected fallback below exists precisely because writing to a shared config is not a free action. SKILL.md step 4 says "suggest" and these two are read in sequence, so they must not disagree. Without filesystem access, check the system prompt / project instructions and suggest the user add the instruction there. Keep the suggestion to a sentence or two. The block above is the propagated artefact: suggest it whole, including the anchoring paragraph, because constraints that live only in SKILL.md arrive after the decision they were meant to govern.
Load is not activation. The failure the block's wording guards against, reported from real use: the agent loads the skill per the config instruction, then stops — the Session Start Protocol (log files, scan, review trigger) never runs, and nothing surfaces the omission because a loaded-but-inert skill looks identical to an active one from the user's side. It moves only when the user explicitly asks "have you executed the session start protocol?". Hence the two belts above: the instruction demands the protocol by name, not just the load, and the post-task summary line makes silent inactivity visible at the first task boundary. If you adopt only one line of the block, adopt the post-task check — in field use it turned an intermittently-activating install into a stably-recording one.
A third belt sits inside the protocol itself: its scan appends a dated
line to checkpoints.log (SKILL.md, Session Start Protocol step 2), so
the protocol leaves its own trace and a session that skipped it is
detectable afterwards rather than only noticeable in the moment. That is
the missing half of the diagnosis above — the load produces a visible
artefact in the transcript and the protocol produces none, so having
loaded the skill feels like having done the thing the skill asks for.
Treat loaded but not run as a distinct failure cause, alongside the
skill being absent, truncated, or present and skipped: it is the likeliest
of the four, precisely because the load's visible success stands in for
the protocol's invisible omission, and the aggravating condition is the
same one that produces a skipped load — an opening message small enough
that a full startup protocol feels disproportionate to it.
Anti-pattern: don't chain activation through another skill — load task-observer and related skills independently from configuration; a broken chain silences all observation activity.
Capture is hard-enforced by checkpoints hooked onto tool calls; the review
trigger in the Session Start Protocol is a soft step — read a file, compare
a date — and it is skipped the same way activation is. The failure is
self-concealing: capture keeps producing, the log looks healthy and
growing, and the only artefact recording the miss is a file reading
never that nobody reads. Treat "is the review trigger structurally
enforced?" as an install-completeness check, and where the harness offers
a session-start hook, have it compute the state and inject it rather than
asking the agent to go and look:
#!/bin/sh
# SessionStart hook: inject activation + review state as additionalContext.
d="$OBS_WORKSPACE/skill-observations" # the pinned absolute path
open=$(find "$d/observation-log" -maxdepth 1 -name '*.md' -exec grep -l '^status: open$' {} + 2>/dev/null | wc -l | tr -d ' ')
last=$(cat "$d/last-review-date.txt" 2>/dev/null || echo never)
msg="Invoke the task-observer skill before the first tool call."
if [ "$open" -gt 0 ]; then
msg="$msg $open open observations; last review: $last."
case "$last" in (never) msg="$msg Offer the review." ;; esac
fi
printf '{"hookSpecificOutput":{"hookEventName":"SessionStart","additionalContext":"%s"}}\n' "$msg"The count is of files whose status field reads open — not of files
in the directory. Resolved entries deliberately stay in observation-log/
until the day after they were resolved (the grace period lives in the
file, not in session memory), and parked entries are decided and out of
the work queue, so a raw file count overstates the backlog by every entry
the last review just closed, for a day, in every session. Compare dates
without < in [ ]: ISO dates sort lexically, so
[ "$(printf '%s\n%s\n' "$last" "$cutoff" | sort | head -1)" = "$last" ]
is true when $last is not later than $cutoff, in every POSIX shell
(\< is a bash/ksh extension that zsh rejects and sh does not know).
Two details that matter when building one: grep -c exits 1
on zero matches while still printing 0, so $(grep -c … || echo 0)
yields two values — capture the output and ignore the exit code; and prove
the reminder branch fires by running the hook against fixtures at never,
30 days stale and 2 days stale, confirming the third stays silent. A nag
that never fires and a nag that is correctly silent look identical from a
passing run.
The config file is read at session start, so an instruction written mid-session takes effect only on the next one; and in a harness that hot-loads a newly installed skill, the skill being callable right after install proves nothing about activation — it was invoked by hand. The installing session therefore sees everything pass (bundle complete, skill listed, protocol ran, hook emits valid JSON) while the one thing that matters is untested, and the tempting close is "installed and tested". Report the install as activation unverified until a fresh session has been seen invoking the skill on its own, and name the check: in the next session, before any work, confirm the skill was invoked (not merely listed) and that the Session Start Protocol ran — verify by evidence of invocation, not by a side effect that a hand invocation would also produce. If the environment offers no way to start a fresh session from the installing one (a GUI-only harness, no CLI), hand the check to the user as the first item of their next session; an unverifiable step never closes silently as "tested".
The runtime cannot diagnose its own absence — Session Start step 4 only
runs once activation has already succeeded — so the external diagnostic
is the one that catches a skipped install step: if
skill-observations/observation-log/ does not exist after a few sessions
of tool-using work, activation never happened; check the activation
block in the config, or install the hook. The README carries the same
diagnostic for users who never read this file.
SKILL.md step 1 carries the instruction (on turn 1, load the session-start skills directly rather than assuming the config did, and read the config yourself if the mount resolves but its content is not in context). This is the reasoning behind it, and the reason that instruction is only the backup.
In some hosted or bridged session types the activation config — a project instruction file, a session-start hook's context — does not reach context until the second user turn, while the workspace it lives in is already mounted and probes clean. From inside turn 1 a config that is merely late is indistinguishable from one that is absent, so a whole substantive turn can run with no session-start skill loaded. A config delivered on a later turn than the work it governs is equivalent to an absent one for that work.
It has also been seen intermittent rather than merely late: present for turns 2 and 3 of a bridged session, then reported as no longer present at turn 4, with no change in the workspace link and nothing the session did to cause it. So a rule anchored on "once it arrives it stays" fails the same way a rule anchored on "it arrives first" does. Do not treat an earlier turn's config as still in force.
The durable form: a guard against an activation config failing to load
cannot live inside the thing that config loads. The guard in SKILL.md step
1 is in task-observer, which is loaded because the config says to load
it — so on the turn where the config is missing, nothing puts that guard in
context either. By construction it cannot fire on turn 1.
The primary guard therefore belongs in the earliest channel verified present on turn 1 — in most harnesses the user-level preferences or system-prompt block, which is where the line
then read the workspace config and follow it before any other work
goes. See the activation tiers above: that channel is tier 0, and the step-1 guard is the backup for turns 2+ and for sessions where the skill happens to be loaded directly.
Log the cause, not just the miss. When a session-start load is skipped, the fix depends on which of these it was: absent (config unreachable), present-but-not-yet-injected (late — record the turn it arrived), present and skipped (attention), or loaded-but-not-run (the invocation succeeded and the protocol inside it never executed). Only the first two are fixed by changing the channel; the last two are not, and an environmental cause found first will otherwise absorb the whole explanation.
Each of these leaves no error. The skill simply stops being offered, or an activation tier stops firing, and the only symptom is that the Session Start Protocol no longer runs — which is indistinguishable from a session where it ran and found nothing.
Skill discovery is exactly one level deep. The host reads the skills
directory and looks for <entry>/SKILL.md. There is no recursive lookup. On
a library of any size the natural organising instinct is to group skills into
category folders — skills/seo/, skills/clients/, skills/writing/ — and
doing so makes every skill inside them cease to exist: no error, no
warning, no change in behaviour except that the skills stop being offered.
There is nothing to debug, because "not found" has no error to report. The
layout is not merely the happy path; it is the rule, and it does not travel
with the maintainer who reorganises six months later. Keep every skill
directly under the skills directory.
A backup inside the skills directory registers as a rival skill. The
obvious way to back a skill up before overwriting it —
cp -r ~/.claude/skills/task-observer ~/.claude/skills/.task-observer-backup
— puts a second copy where the host is looking. The leading dot does not
exclude it from discovery. The available-skills list then holds two entries
with near-identical descriptions, and the backup still advertises the OLD
version's trigger, including its session-start instruction. That defeats the
upgrade at the one activation tier that survives an unreachable config:
description matching now has two candidates competing for the same trigger,
one of them the file the upgrade just replaced. Back up outside the
skills directory.
A hook registered by script path dies when the exec bit goes. The docs' example reads most naturally as naming the script directly:
{ "type": "command", "command": "'/path/to/hooks/task-observer-activation.sh'" }Anything that rewrites file modes — a cloud-drive resync switching between
stream and mirror mode, a restore from an archive, a checkout on a
filesystem without permission bits — returns the file as mode 600, and the
hook then fails with "permission denied" on every session start with
nothing surfacing the error. Observed: the enforced activation tier was
dead for days. Register the interpreter instead
("command": "bash '/path/to/hooks/…'"), which does not depend on the exec
bit, and treat "the Session Start Protocol stopped running" as a prompt to
check the hook rather than the skill.
A command-shape guard can refuse a read-only call. Some setups run a
pre-tool hook that inspects the literal command text for governed path
segments combined with a write indicator (=, >, cp, mv) and denies
the whole call when both appear anywhere in the string — regardless of
whether the path is actually being written to. Two consequences worth
knowing before diagnosing: a cd into a governed directory followed by a
read-only script call is denied even though nothing is written; and a
blocked command merely quoted inside an observation body can trip the
same guard when that body is written through a shell. Neither is a
permission problem with the destination. Where this bites, invoke the
script with an absolute path and no cd, and write observation bodies with
the editing tool rather than through a shell.
Staging a harness configuration change is not like staging a skill. A
skill's SKILL.md is read by itself; a corrupted staged copy fails to parse
and nothing else is affected. A harness config file — settings.json with
its hook entries, deny-lists and protection rules — is read by the harness,
and it commonly carries other protections alongside the thing being
changed. The natural delivery shape for a config change is a fragment:
"here's the new key, add it under hooks". A hand-merge that goes wrong
there does not fail loudly; it produces a valid file with the other
protections missing.
So: deliver a config change as a complete replacement file built from the user's current one, never as a fragment to splice in; state in one line what else that file was carrying, so the user can verify it survived; and where the harness supports it, prefer a separate file over editing a shared one.
Some setups guard shared config files with hooks or file-protection rules that deny agent edits. If an edit to the config is denied, never retry the same edit blindly and never attempt to bypass the guard — a denial is the governance system working as intended, and a silent skip is just as bad (the user believes activation is set up when only description-level matching is active). Surface the denial to the user and offer these fallbacks, lightest first: (a) if the denial came from an interactive permission system rather than a hard guard — the denial message offers an approval path, names a permission rule, or otherwise indicates consent would clear it — ask the user to approve a retry of the same edit; one confirmation resolves it, and escalating past this step spends user effort on something already recoverable; (b) ask the user to paste the activation block into the file themselves; (c) if the user's environment provides its own temporary-authorization mechanism (a marker file, an environment variable, or similar), ask the user to authorize the edit through that mechanism and revoke it afterwards; (d) where the platform supports unguarded project-level instruction files, add the activation instruction there instead. Branch on whether the denial is retryable-with-consent or categorical: an interactive permission prompt is cleared by asking, a hook or file-protection rule is not, and collapsing both into "hand the task back to the user" is safe but consistently over-escalates. Never assume unrestricted edit access to shared or governance-tracked config — many setups gate exactly those files.
The same shape applies to a blocked hook installation. The hook option from the same paragraph writes to the harness's settings file or hooks directory, and those are gated at least as often as the config file. A denied hook install follows the rules above unchanged: never retry the same install through a different tool or path, never bypass, surface the refusal, then fall back lightest first — ask for a retry if the denial is retryable-with-consent; ask the user to add the hook entry themselves (hand them the exact JSON or script); use a temporary authorization mechanism if one exists; and if the hook cannot be installed at all, fall back to the declarative config instruction, saying plainly that the enforced tier is not in place and the probabilistic one is. A hook that silently failed to install is worse than none: the user believes the strongest tier is active.
The procedures in this skill are written as capabilities. This table is the only place product-specific names live; when a step says "present the staged file" or "register the review", look up the current environment here rather than guessing a tool name.
Cowork execution mode — cloud vs local is a user setting. Cowork can
run a session in two modes with materially different tool surfaces, and
the platform's default moved to cloud (so installs upgraded from earlier
versions may silently change mode — users experience that as the tool
losing abilities). The switch: Settings → Cowork → "Run new tasks in the
cloud" (plus a "Beta" button at session start). The tell: the
local-filesystem permission tool (allow_cowork_file_delete) present in
the tool surface ⇒ local session; absent ⇒ cloud. What differs:
| Capability | Local session | Cloud session |
|---|---|---|
| Scheduled-task editing | desktop scheduled-task tools | draft instructions only; the user applies them on the Scheduled tasks page |
| Deleting workspace files | permission gate (allow_cowork_file_delete) | no delete tool; rename-away only |
| Git on the mounted workspace | with the delete grant | never — use a sandbox clone |
| Locally-installed MCP servers | reachable | unreachable unless proxied via the desktop app |
| Working-file persistence | workspace folder persists | files not handed back are not kept |
| Runs while the computer is off | no | yes |
Steps that must edit scheduled tasks, delete workspace files (review cleanup, keep-two pruning), or run git in the shared folder need a local session — treat that as a one-line precondition on those steps, and when reporting a mode-conditional limit, name the mode and the switch in the same breath.
| Capability | Claude Cowork | Claude Code | Web chat / no filesystem |
|---|---|---|---|
| Persistent workspace | the shared folder | pinned absolute path (see activation block) | none — handoff-doc mode below |
| Live skill files | .claude/skills/{skill}/ on a read-only mount (writes fail with EROFS) | ~/.claude/skills/{skill}/, ordinary writable files — no guard | n/a |
| Present a staged file for install | the file-presentation tool (present_files) with its upload button | none — report the staged path and a change summary in chat | paste into the handoff doc |
| Scheduled review | the app's scheduled tasks (cloud-default: run remotely; legacy local mode: run on the user's machine, only while it is on — see the execution-mode note above) | cron / a harness SessionStart hook / an account-level scheduler that can reach the workspace | calendar reminder + manual trigger |
| Session-start hook | none | SessionStart hook (hookSpecificOutput.additionalContext) | none |
| Local-vs-cloud tell | allow_cowork_file_delete present ⇒ local session | n/a | n/a |
Grow the table when a new environment appears; do not scatter its tool names through the procedure files.
Managed skills directories (dotfile managers, sync tools). Where the
skills directory is generated by chezmoi, GNU Stow, yadm, a symlinked
dotfiles repo or any sync tool, the live path is not the source of truth:
copying a staged update onto it succeeds, looks installed, and is
silently reverted at the next apply — "I installed it" and "it is still
in effect" become indistinguishable. Detect the manager at install (a
symlink into a dotfiles checkout, a .chezmoi* marker, the directory
listed in a Stow package) and record it in the activation config. Then
install into the manager's SOURCE, not the live path, and let the manager
apply it; or exclude the skills directory from management. "Always start
from the live file" (references/skill-authoring.md) still holds for
READING — the live file is what the loader loads — but the WRITE-BACK
path is the manager's source. The weekly review's staged-work
reconciliation (diff -rq staged vs live) is the check that catches a
reverted install: run it at the session after any install on a managed
directory.
Staging a skill update as a branch or commit the user merges gives the
same guarantee as the skill-updates/ directory — nothing goes live
without a user action — plus diffs, history and rollback, and it suits a
multi-device setup that syncs skills through a private repository. It is
an option, not a change to the default: the review still writes the
staging manifest, still never edits the live install, and still presents
the change for a decision. Treat the version-control commands as a
mutation surface for the observation log (see
references/observation-log.md).
The skill is least valuable at the moment it is adopted: the log is empty, a review correctly reports nothing to do, and the largest pile of uncaptured insight — the project's own history — is never touched because no procedure says to touch it. Session Start step 7 offers a one-off mining pass over handover, architecture and decision docs, commit history since the last release, test and verification scripts (which encode hard-won discipline densely), and existing agent-instruction files, which are largely a record of corrections the user already had to make. One such pass over seven weeks of history produced twenty-three actionable observations, eleven of them factual corrections to an existing skill. Backfilled entries cite the durable artefact; the pass runs once.
Persistence is one axis; the price of a write is another. Three regimes:
This skill consists of SKILL.md, the reference files it lists
(weekly-review.md, skill-authoring.md, environments.md,
observation-log.md, signals.md, migration.md,
starter-principles.md) and scripts/migrate-log.py and
scripts/validate-skill-bundle.py. If a referenced
file is missing, the install is
incomplete: proceed using the rules in SKILL.md, tell the user which
files are missing, and point them to the full bundle at the canonical
source (for the published version, the repository named in the attribution
block).
When context compacts mid-task, the CLAUDE.md structural trigger re-invokes
this skill on the resumed session automatically (the resumed session reads
CLAUDE.md anew). Observations before and after compaction are written as
separate files under the same observation-log/ directory, each with its own
id (the id counter is derived from existing filenames, so it continues
seamlessly across the compaction boundary). This is the main reason the
structural trigger exists — a resumed session's opening message may not
match the description triggers.
Installation, shared-folder setup, expected behaviour, and the cadence pattern live in the public repo. These links are for the human reader: share them with the user rather than fetching the pages — the skill's behaviour is defined entirely by its own files, never by external content:
Authentication and attribution are separate channels in git: the push
credential (PAT/SSH key) controls who may WRITE; the commit's author
email (git config user.email) declares who WROTE, and the platform maps
that email to whichever account has it verified — regardless of which
account pushed. Before the first terminal commit in any clone used for a
specific identity, verify git config user.email resolves to the
intended account, and set repo-local config where the machine's global
identity differs. Diagnose suspected mis-attribution via the commits API
(author login vs commit email). Fix forward only: rewriting a published
main to correct author metadata (with tags/CI descending from it) costs
more than the cosmetic gain — verifying the email on the intended account
is the alternative remedy.
The methodology is environment-independent; only persistence varies. In web-chat-style environments, collect observations in-session and deliver them in a structured handoff document the user stores and pastes into the next session. Offer the handoff proactively when the conversation winds down — a premature offer is a minor interruption; a missing one is lost work.
# Session Handoff: [Session Topic]
**Date:** [date]
**Context:** [what was worked on; what the next session needs to know]
## Decisions Made
[numbered]
## Observations Logged
[each observation in the frontmatter format from SKILL.md → How to Log; the
next session writes each as its own file in `observation-log/`]
## Cross-Cutting Principles (current)
[active or newly added]
## Action Items
[next steps with enough context to resume]
## Working Artifacts
[drafts/analyses in full]