Runs a command and reads the telemetry that run just emitted, returning a `verify-behavior` receipt (confirms / contradicts / ambiguous / null) that grades a behavioral assertion — span count, parent/child structure, duration, error-path status, fan-out, attribute cardinality, ordering — against the observed spans, never against the diff read back. Inputs are a command plus an expectation set. Walks two cheapest-first rungs (an in-memory/file exporter, then `dash0 -X otlp proxy --agent-mode` plus `dash0 spans query`), stamps two-layer run identity, and self-skips when the repo has no Observability Profile dev target. Vendor-neutral: OTLP is the contract, Dash0 one implementation of the read. Use when a fan-out count, a retry that fired twice, an error swallowed into a 200, an N+1, or an unbounded cardinality needs proof from a real run rather than a reading of the code. Triggers only on an explicit ask — "observe this run's telemetry", "prove this behavior with a trace", "/observe-run".
64
80%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Low
Low-risk findings worth noting
Fix and improve this skill with Tessl
tessl review fix ./skills/quality/observe-run/SKILL.mdRun a command, then read the telemetry that run just emitted, and grade a stated behavioral claim against it.
A test asserts through a keyhole. A trace of the same execution exposes the whole thing: fan-out that went from 2 calls to 14, a retry that fired 5 times, an error swallowed into a 200, an N+1 query, a user ID landing as unbounded metric cardinality — none of which turns a test red. This skill closes that gap by giving the dev loop a fourth telemetry role this repo did not have: telemetry as verification evidence, read seconds after the run that produced it, rather than telemetry read as coverage, as yesterday's production exposure, or as a post-deploy confirmation.
This
SKILL.mdis a thin index. Detailed rules live inrules/*.mdand load on demand.
The inputs are a command plus an expectation set — nothing else.
Skill("observe-run")
command: <the command that runs the code under test> # required
expectations: <one or more behavioral assertions> # required; by-construction ones are refused
rung: 1 | 2 # optional; default resolved by rules/rungs.md
run_id: <string> # optional; defaults to a generated dev.run.id
caller: <invoking skill or agent> # logging onlyReturn value is exactly one verify-behavior receipt, whose final line is
[receipt] verdict: <confirms|contradicts|ambiguous|null> — or nothing at all when the
self-skip gate below fires.
This skill has zero dependency on any aw artifact. It takes no worktree for granted, reads no
plan artifact, and checks no acceptance-criteria ledger. It runs anywhere a command can run.
Before doing any work, resolve the repo's committed Observability Profile
(<repo>/memory/observability-profile/INDEX.md, written by measurable setup — see
skills/quality/measurable/rules/setup-profile.md) and
check it names a dev run target (the field measurable setup now asks for — see
Integration: measurable).
| State | Action |
|---|---|
| No profile at all | Self-skip. No tokens spent, no report line emitted — not "ran and reported skipped." |
| A profile exists but names no dev run target | Same self-skip. If this was an explicit /observe-run invocation (not a composed call from another skill), also emit one line pointing at Skill("measurable", "setup") — an explicit human ask deserves an answer; a composed call stays silent. |
| A profile exists and names a dev run target | Proceed to Workflow. |
This mirrors the quiet-exit shape of the pr-reviewer measurability lens
(agents/shared/rules/measurability-review.md):
a clean state produces nothing, not a line saying so.
The expectation set is checked too: an empty expectation set has nothing to assert, so this
skill returns the same quiet self-skip rather than a vacuous confirms.
| Step | Name | Rule file | Gate |
|---|---|---|---|
| 1 | Self-skip gate | (this file, above) | Dev run target resolvable, expectation set non-empty |
| 2 | Assertion provenance | rules/assertion-provenance.md | Every expectation is behavioral, never by-construction |
| 3 | Read dev-run lessons | rules/lessons.md | Merged repo+global lessons bias the run; a missing LoreKit costs nothing |
| 4 | Rung selection | rules/rungs.md | Cheapest rung that can decide the claim; rung 1 default |
| 5 | Run identity | rules/run-identity.md | Dataset + the three resource attributes stamped |
| 6 | Read | rules/reader-adapters.md | Reader resolved; a missing Dash0 CLI costs rung 1 nothing |
| 7 | Verdict | rules/receipt-mapping.md | Proxy-stream state (or rung-1 span list) mapped to one of the four canonical tokens |
| 8 | Write dev-run lesson | rules/lessons.md | On friction only; never blocks the receipt; drops secrets |
rules/assertion-provenance.md's discriminator. Refuse any
expectation satisfiable by reading the source alone, and name the behavioral alternative — never
run it, never grade it.repo:: + global lessons in the
loop::observe-run-lessons bucket and apply the matches as considerations that bias the rung,
the run-identity mechanism, and the reader choice. A missing LoreKit connection makes this a
silent no-op. See rules/lessons.md.dash0 spans query) only for cross-process fan-out, multi-service
claims, or a baseline comparison. See rules/rungs.md.dev.run.id; stamp it alongside
deployment.environment.name and vcs.ref.head.name. See
rules/run-identity.md.rules/reader-adapters.md.confirms / contradicts / ambiguous / null, per
rules/receipt-mapping.md, and return the receipt.rules/lessons.md.| Concern | Owner |
|---|---|
| The four canonical receipt verdicts and their shape | verify-behavior — rules/receipt.md. This skill maps into it and never redefines it. |
| The cheapest-first ladder vocabulary (Tier 1/2/3) | verify-behavior — rules/ladder.md. This skill's two rungs live inside that ladder's Tier 3. |
| The by-construction framing | test-provenance-guard — cited by |
rules/assertion-provenance.md, never forked. | |
| The self-improvement lesson loop (bucket, schema, recurrence, promotion gate) | the external lorekit-setup skill (rules/self-improvement-loops.md). rules/lessons.md instantiates it for dev-run lessons; it never redefines it. |
measurable implementmeasurable's implement mode calls this skill as a prove-it step after writing instrumentation,
turning a static file:line claim into an executed one. Skipped with one report line when this skill (or its
prerequisite dev run target) is unavailable — advisory, consistent with measurable's own Core
Principle 6.
verify-behavior Tier 3verify-behavior/rules/ladder.md's Tier 3 table gains a third approach delegating here, alongside
"run the covering test" and "synthesize a minimal repro." Skipped with one report line when this skill is not
installed.
fix-bugFor a telemetry-sourced bug, fix-bug's Phase 2.5 reproduction step checks repro fidelity, and the
division of labour is the important half: fix-bug reads the originating production span's shape
out of its own Evidence Record and passes it in as literal expectations; this skill then grades the
local run against them.
observe-run never fetches the production span and cannot. Both rungs read only the telemetry the
run they just executed emitted, scoped to that run's dev.run.id in the dev dataset
(rules/run-identity.md) — so an expectation phrased as "matches the
production span" asks this skill to assert on data its reader is structurally unable to see, and it
would return null every time.
aw coupling. No autonomous-workflow workspace artifact, no plan document, no worktree
assumption. This skill runs anywhere a command runs.confirms / contradicts / ambiguous / null —
never a fifth token.rules/lessons.md) is additive over the Observability Profile, never a
relocation of it: the profile stays committed and repo-scoped so it is diffable and
teammate-visible; lessons live in LoreKit (global + repo::) so they auto-update and merge at
read. A missing LoreKit connection costs the loop, and the run, nothing.X exists" instead of a behavioral claim about what the run did.dash0.cli.otlp_proxy.forwarded delivery-receipt event as proof of span content —
content comes only from dash0 spans query.error event, *.failed > 0, a deadline-hit shutdown) as
contradicts — it is ambiguous, never proof of absence.dev.run.id as a metric dimension.rules/lessons.md).confirms/contradicts/ambiguous/null is the
receipt's job; a lesson is only about running and reading better next time.rules/run-identity.md scopes it to — dev.run.id on spans and logs only, never metrics.
At rung 2 that attribute is a requirement, not a best effort: a run that cannot stamp it
falls back to rung 1 rather than proceeding without it.rules/receipt-mapping.md.repo:: + global) before rung selection, and a lesson
was written only on genuine friction, at the scope its classification picks, with any
secret-bearing candidate dropped. A missing LoreKit connection left both steps silent and did
not block the receipt.39b3f44
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.