CtrlK
BlogDocsLog inGet started
Tessl Logo

measurable

Ensures every delivery ships with the telemetry needed to prove its impact and catch its own regressions: RUM/analytics events for user-facing web changes (delegates to rum-tracking), OpenTelemetry traces, metrics, and structured logs for new or changed API endpoints, and explicit error/warning signal paths so failures surface instead of going silent. Knows OpenTelemetry Weaver — validating a signal against a semantic-convention registry, and treating a renamed metric or attribute as the breaking change it is. Modes: guide (default), implement, audit, setup. `setup` interviews the project once, recording its telemetry stack, schema (Weaver registry), per-package instrumentation approach for monorepos, and regression-detection expectations as a committed Observability Profile. Triggers on "is this measurable", "add telemetry", "instrument this endpoint", "check observability coverage", "will we know if this regresses", "does this break the telemetry contract", "/measurable".

68

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Excellent workflow design and actionable process guidance, organized as a thin index with per-mode required reading. The critical defect is that every referenced rule and template file is absent from the bundle, leaving the skill's operational payload unverifiable; secondary costs come from repeated policy statements across sections.

Suggestions

Ship the referenced rules/*.md and templates/*.md files in the bundle (or inline the decision-critical content, e.g. the audit checklist and regression-signal criteria) so the thin index points at real files.

State the companion/degradation policy once (e.g. in Core Principle 6) and reference it from the intro note, implement Step 6, audit Step 4, and the related anti-patterns instead of restating it in each place.

Trim implement Step 6's meta-explanation of observe-run's provenance rule to the decision-relevant rule (phrase expectations from the closed list; avoid by-construction assertions) and move the rationale into the observe-run reference.

DimensionReasoningScore

Conciseness

The body is dense and skill-specific — no space is spent explaining concepts Claude already knows, and sections like the mode-detection table and one-liner anti-patterns are token-efficient. It is not 5 because the advisory/degradation policy ('never block, skipped with one report line') and the 'rename is a breaking change' rationale each recur across four to six sections, and implement Step 6's meta-explanation of observe-run's provenance rule could be trimmed to the decision-relevant parts.

4 / 5

Actionability

Concrete executable guidance throughout: 'Skill("rum-tracking", "implement", "<target>")', re-running 'weaver registry check' in the same diff, worked expectation phrasings for observe-run, the missing/unlinked/pass finding taxonomy, and file:line citation requirements. It is not 5 because the code-level instrumentation detail is deferred to rule files and no instrumentation code appears inline; it is above 3 because what is inline is directly executable rather than pseudocode.

4 / 5

Workflow Clarity

Four modes each carry a numbered, sequenced workflow with explicit validation checkpoints: implement Step 6 ('Prove it') converts static claims into run-verified expectations, the registry check is re-run in the same diff, audit classifies findings into three verdicts with mandatory file:line citations, and setup is idempotent. It is not 4 because checkpoints are explicit and include feedback loops (fix-and-revalidate, advisory-to-strict escalation) rather than being implicit.

5 / 5

Progressive Disclosure

The structure is exemplary — a self-declared 'thin index' with one-level-deep, clearly signaled references and a per-mode required-reading table — but none of the seven referenced files (rules/scope-detection.md, rules/frontend-rum.md, rules/backend-instrumentation.md, rules/regression-signals.md, rules/weaver-schema.md, rules/audit-checklist.md, rules/setup-profile.md, templates/observability-profile.template.md) exist in the bundle. As shipped, it is an index whose payload is missing, so the reference chain cannot be verified; it is above 2 because the inlined content is appropriate for an index and references are clearly signaled rather than buried.

3 / 5

Total

16

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete capability inventory across frontend, backend, schema, and setup concerns, and an explicit trigger list of natural user phrases. The only gap is a few missing common synonyms (e.g., 'monitoring', 'alerting') in the trigger terms.

DimensionReasoningScore

Specificity

Lists multiple specific concrete actions with comprehensive coverage: 'RUM/analytics events for user-facing web changes (delegates to rum-tracking), OpenTelemetry traces, metrics, and structured logs for new or changed API endpoints, and explicit error/warning signal paths', plus 'validating a signal against a semantic-convention registry' and a concrete description of the setup interview. It is not 4 because frontend, backend, schema, and setup capabilities are all named with specific technologies, leaving no material coverage gap.

5 / 5

Completeness

Explicitly answers both questions: the 'what' is a detailed capability inventory (RUM delegation, OTel traces/metrics/logs, error/warning signal paths, Weaver registry validation, setup interview) and the 'when' is an explicit 'Triggers on' clause with concrete phrases. It is not 4 because the when-guidance is not merely present — it is explicit and specific with natural trigger phrases.

5 / 5

Trigger Term Quality

Explicit trigger list covers natural phrases users would actually say: 'is this measurable', 'add telemetry', 'instrument this endpoint', 'check observability coverage', 'will we know if this regresses', 'does this break the telemetry contract', '/measurable'. It is not 5 because common synonyms such as 'monitoring', 'add logging', or 'alerting' are missing; it is well above 3 because the phrases are varied, natural, and include a slash command.

4 / 5

Distinctiveness Conflict Risk

Clear niche (telemetry/observability gating with named tooling: Weaver, OpenTelemetry, rum-tracking) with explicitly stated boundaries against the nearest overlapping skills ('delegates to rum-tracking', otel skills 'invoked via Skill()'). It is not 4 because the description itself draws the distinction from the closest neighbors rather than leaving minor overlap risk.

5 / 5

Total

19

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 25 missing, 3 suspicious

Warning

Total

13

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.