CtrlK
BlogDocsLog inGet started
Tessl Logo

observability-and-instrumentation

Instruments code so production behavior is visible and diagnosable. Use when adding logging, metrics, tracing, or alerting. Use when shipping any feature that runs in production and you need evidence it works. Use when production issues are reported but you can't tell what happened from the available data.

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An unusually actionable, well-sequenced instrumentation guide whose code examples and validation steps are exemplary. Its two weaknesses are verbosity in the prose-heavy entry-point and rationalization sections, and a single external reference that resolves to a missing file.

Suggestions

Fix or remove the reference to `../../references/observability-checklist.md` — no references/ directory exists in the bundle, so the link dead-ends.

Trim the entry-point attribution prose (currently ~15 lines) to the rule plus the code example, and cut the rationalizations table to the 3–4 rows that aren't already covered by Red Flags.

Consider moving the runbook-writing subsection and the full verification checklist into the (real) reference file to shorten SKILL.md toward an overview.

DimensionReasoningScore

Conciseness

Nearly every section delivers non-obvious judgment (RED/USE, cardinality rules, symptom-vs-cause alerting, the entry-point attribution rule) rather than explaining concepts Claude already knows. Not a 5: passages like the entry-point prose ("A field that merely correlates with an entry point is a hint, not an attribution...") and the seven-row rationalizations table run long and could be tightened; not a 3: the padding is minor relative to the density of the whole.

4 / 5

Actionability

Copy-paste-ready TypeScript for structured logging, the Express correlation-ID middleware, the `runLog` helper, a prom-client histogram with explicit buckets/labels, and the OpenTelemetry SDK setup, plus a filled-in runbook template and concrete alert rules. Not a 4: examples are complete and executable and cover the common cases across all three signal types.

5 / 5

Workflow Clarity

A clearly sequenced 7-step process that starts from defined questions, and step 7 ("Verify the telemetry itself") is explicit validation with concrete checks, reinforced by a nine-item final Verification checklist. Not a 4: validation isn't just present, it's a dedicated step with feedback loops (force an error in staging, find it by requestId, fix and re-fire alerts).

5 / 5

Progressive Disclosure

Sections are well organized, but the only external reference — "see `../../references/observability-checklist.md`" — points to a file that does not exist in this bundle (no references/ directory), so navigation dead-ends. Not a 4: a broken reference is more than a minor organization gap; not a 2: structure is real and the reference is clearly signaled at one level deep.

3 / 5

Total

17

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with explicit, multi-scenario 'Use when' triggers and natural on-call vocabulary. Its main weakness is that the 'what' is a single abstract capability statement rather than a list of concrete actions, and it lacks synonyms like 'monitoring' or 'telemetry'.

Suggestions

Replace the abstract capability statement with concrete actions, e.g., "Adds structured logging, RED metrics, distributed tracing, and symptom-based alerts to production code."

Add the common synonyms users actually say — "monitoring", "telemetry", "observability" — to the trigger clause.

Disambiguate from adjacent skills in the description itself (e.g., "for diagnosing a live failure, use debugging") rather than only in the body.

DimensionReasoningScore

Specificity

The 'what' is one abstract action — "Instruments code so production behavior is visible and diagnosable" — which names the domain but states no concrete instrumentation actions (e.g., add structured logs, define RED metrics, wire up traces). Not a 4: it does not list several specific actions; not a 2: the domain is precise and the signal types (logging, metrics, tracing, alerting) do appear, though only as trigger terms.

3 / 5

Completeness

Explicitly answers both parts: what ("Instruments code so production behavior is visible and diagnosable") and when, via three concrete "Use when..." clauses covering feature work, incident response, and shipping. Not a 4: the 'when' is fully explicit and multi-scenario rather than only somewhat specific.

5 / 5

Trigger Term Quality

Natural phrases users would say are well covered: "adding logging, metrics, tracing, or alerting", "production issues are reported", "you can't tell what happened from the available data", "evidence it works". Not a 5: common synonyms like "monitoring", "telemetry", "dashboards", or "observability" itself are absent; not a 3: coverage goes beyond a couple of generic keywords with three distinct trigger scenarios.

4 / 5

Distinctiveness Conflict Risk

The observability/instrumentation niche is clear and triggers like "adding logging, metrics, tracing, or alerting" are distinct. Not a 5: "Use when production issues are reported but you can't tell what happened" overlaps with a debugging/diagnosis skill (the disambiguation lives in the body's 'NOT for' section, not the description); not a 3: the overlap risk is limited to closely adjacent skills.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
addyosmani/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.