CtrlK
BlogDocsLog inGet started
Tessl Logo

incident-triage-runbook

The SRE team's runbook for triaging production latency and error-rate incidents. Use this whenever investigating an incident, a latency spike, elevated error rates, or when asked "what caused X" about a production service.

69

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally tight, expert-level runbook that sequences the investigation clearly and bakes in hard-won heuristics. The minor gaps are literal command examples and an explicit validation feedback loop, neither of which is essential for a read-only triage workflow.

DimensionReasoningScore

Conciseness

Lean imperative prose that assumes Claude's competence ('Don't open the log first', 'Don't grep to fish') with zero concept padding; every line earns its place.

5 / 5

Actionability

Provides concrete named metrics (p99_latency_ms, error_rate, db_pool_utilization) and services (checkout/cart/auth/inventory) with a precise procedure, but stops short of literal copy-paste CLI/query commands, leaving minor execution gaps.

4 / 5

Workflow Clarity

A clear numbered 1-5 sequence with conditional branches ('If a deploy lines up' / 'No deploy lines up ->') and a confirm step, but it lacks an explicit validate->fix->retry feedback loop (less critical for read-only triage, but absent nonetheless).

4 / 5

Progressive Disclosure

A self-contained skill well under 50 lines with no need for external references and three cleanly organized sections, which the rubric explicitly allows to score 5 for simple skills.

5 / 5

Total

18

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-targeted description that clearly states both purpose and trigger conditions with natural SRE language. Its only soft spot is specificity: 'triaging' is one umbrella action rather than a list of concrete operations.

DimensionReasoningScore

Specificity

Names a concrete domain (production latency/error-rate incidents) and one core action (triaging), but does not enumerate several specific sub-actions, matching the 'domain + 1-2 concrete actions' anchor rather than the 'several specific actions' anchor above.

3 / 5

Completeness

Explicitly answers both what ('the SRE team's runbook for triaging production latency and error-rate incidents') and when ('Use this whenever investigating an incident, a latency spike, elevated error rates, or when asked what caused X') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Includes natural SRE phrasing with synonyms — 'investigating an incident', 'a latency spike', 'elevated error rates', 'what caused X' — but omits common variations like alert, page, 5xx, or SLO burn, so it is strong but not fully comprehensive.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (SRE production incident triage) with distinct, domain-specific triggers and minimal realistic overlap with other skills.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
anthropics/cwc-workshops
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.