CtrlK
BlogDocsLog inGet started
Tessl Logo

observe-run

Runs a command and reads the telemetry that run just emitted, returning a `verify-behavior` receipt (confirms / contradicts / ambiguous / null) that grades a behavioral assertion — span count, parent/child structure, duration, error-path status, fan-out, attribute cardinality, ordering — against the observed spans, never against the diff read back. Inputs are a command plus an expectation set. Walks two cheapest-first rungs (an in-memory/file exporter, then `dash0 -X otlp proxy --agent-mode` plus `dash0 spans query`), stamps two-layer run identity, and self-skips when the repo has no Observability Profile dev target. Vendor-neutral: OTLP is the contract, Dash0 one implementation of the read. Use when a fan-out count, a retry that fired twice, an error swallowed into a 200, an N+1, or an unbounded cardinality needs proof from a real run rather than a reading of the code. Triggers only on an explicit ask — "observe this run's telemetry", "prove this behavior with a trace", "/observe-run".

64

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./skills/quality/observe-run/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

62%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body excels at workflow orchestration — gates, fallbacks, and a completion checklist — and is free of filler explanations of known concepts. Its weaknesses are heavy internal repetition and a thin-index design whose referenced rule files are absent from the bundle, leaving the executable specifics unreachable.

Suggestions

Consolidate the duplicated content: drop either the workflow table or the step-by-step prose, and state the four verdict tokens once, referencing them thereafter, to reclaim a substantial token budget.

Ship or inline minimal executable specifics — a rung-1 in-memory/file exporter setup command and a sample `dash0 spans query` invocation — so the skill remains actionable even when the rules/*.md files are not present.

Trim the three cross-skill integration sections to one line each (skill name + delegation contract), moving the division-of-labor detail into the referenced skills' own files.

DimensionReasoningScore

Conciseness

The body repeats itself well beyond minor trimming: the four verdict tokens are enumerated roughly five times, the self-skip rule is restated in the gate section, step 1, Core Principle 4, and the Definition of Done, and the workflow appears as both a table and a near-duplicative step-by-step prose walkthrough. It is dense with skill-specific (not generally known) content, so it is above the noticeably-verbose anchor 2, but it clearly could be tightened, matching anchor 3.

3 / 5

Actionability

Concrete elements exist — the input schema with required/optional fields, the receipt's final-line format, the profile path, and the three resource attributes — but every executable mechanic (rung-1 exporter setup, the dash0 query invocation, the provenance discriminator) is deferred to rules/*.md files that are not present in the bundle. This matches the anchor for some concrete guidance with key details missing rather than anchor 4's mostly-executable guidance.

3 / 5

Workflow Clarity

Eight steps are clearly sequenced in a table with an explicit gate per step and then expanded in prose; the self-skip gate validates preconditions before any spend, rung-2 stamping failure falls back to rung 1 rather than proceeding, and a Definition of Done checklist closes the loop. This matches the top anchor: explicit validation steps, an error-recovery path, and a checklist.

5 / 5

Progressive Disclosure

The index structure is exemplary — one rule file per workflow step, each clearly signaled with its purpose and one level deep — but scored against the actual bundle, none of the referenced files (rules/*.md, ../verify-behavior/rules/*, ../measurable/rules/*, ../../../agents/...) exist in this skill's directory. A self-declared "thin index" whose referenced materials do not ship leaves the disclosure chain broken, which sits below the level-4 anchor's "minor organization gaps".

3 / 5

Total

14

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An unusually strong description: it states concrete capabilities, the exact receipt contract, two explicit rungs, vendor neutrality, and precise trigger guidance including the explicit-ask-only constraint. The only weakness is a trigger list that favors constructed phrases over the full range of natural phrasings a user might use.

Suggestions

Add one or two more colloquial trigger phrasings (e.g., "check the traces from this run", "verify with OpenTelemetry") so natural-language variants outside the constructed phrases also invoke the skill.

Trim internal jargon such as "two-layer run identity" and "Observability Profile dev target" to a short gloss, since these terms are opaque to a first-time reader deciding whether to invoke the skill.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions with comprehensive coverage: "Runs a command and reads the telemetry that run just emitted", "grades a behavioral assertion — span count, parent/child structure, duration, error-path status, fan-out, attribute cardinality, ordering", plus rung walking, run-identity stamping, and the self-skip behavior. It matches the anchor for multiple specific concrete actions with no minor gaps, so it is not the level-4 example.

5 / 5

Completeness

It explicitly answers both "what" (runs a command, reads that run's telemetry, returns a verify-behavior receipt grading expectations) and "when" ("Use when a fan-out count … needs proof from a real run rather than a reading of the code" and "Triggers only on an explicit ask") with concrete trigger phrases — a clear match to the top anchor, well above the level-4 anchor where 'when' is only loosely specific.

5 / 5

Trigger Term Quality

Explicit trigger phrases are present ("observe this run's telemetry", "prove this behavior with a trace", "/observe-run") alongside natural scenario terms ("a fan-out count", "a retry that fired twice", "an N+1", "unbounded cardinality"), giving good keyword coverage. It sits below the level-5 anchor because the trigger list leans on somewhat constructed phrasing and omits natural variants a user might say, such as "check the traces" or an OpenTelemetry-mention trigger.

4 / 5

Distinctiveness Conflict Risk

"Triggers only on an explicit ask" combined with a tightly scoped niche — telemetry emitted by a run it just executed, "never against the diff read back" — gives it a clear niche with distinct triggers and minimal conflict risk with other testing or tracing skills.

5 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 20 missing, 4 suspicious

Warning

Total

13

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.