CtrlK
BlogDocsLog inGet started
Tessl Logo

incident-response

Manages production incident response workflows from initial triage through postmortem. Guides responders through alert acknowledgment, mitigation, status communication, and follow-up. Use when the user mentions outages, service degradation, on-call pages, incident response, postmortems, error rate spikes, latency issues, or on-call escalation. Covers rollback decisions, stakeholder updates, and structured post-incident review.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a clear, well-structured incident workflow with concrete time bounds, naming conventions, validation checkpoints, and ready-to-use templates and checklists. Weaknesses are concentrated in Step 4's vague 'do stuff' phrasing and the absence of guidance for mitigation failure paths, plus minor unspecified steps in the mitigation workflow.

Suggestions

Replace 'do stuff to follow up. Things like writing a postmortem, scheduling a review, etc.' with the concrete actions already implied (e.g., 'Write the postmortem, schedule the review meeting, and file action items') — this improves both conciseness and actionability.

Add an explicit fallback for when rollback fails or no recent change is identifiable (e.g., 'If rollback does not restore baseline or no recent change exists, escalate by paging the domain owner and widen the incident channel') to close the workflow-clarity feedback-loop gap.

Make the mitigation step's investigation concrete: specify where to look for the most recent change (deploy dashboard, feature-flag console, config-change log) so the 'identify the most recent change' instruction is executable rather than a hint.

DimensionReasoningScore

Conciseness

The body is largely lean — time bounds ('within 5 minutes', 'every 15 minutes'), a naming convention, and compact checklists — but the Step 4 intro 'do stuff to follow up. Things like writing a postmortem, scheduling a review, etc.' is vague filler that adds tokens without information. It is above the 3 anchor because padding is isolated to one or two lines, not several sections.

4 / 5

Actionability

Most guidance is executable and specific: 'Acknowledge the alert in the on-call tool within 5 minutes', 'Open an incident channel (e.g., `#inc-<short-name>`)', 'Post status updates every 15 minutes', plus a concrete postmortem outline and review checklist. It falls short of 5 because of the 'do stuff to follow up' phrasing and the un-specified 'identify the most recent change' step (no command or where to look), leaving minor gaps.

4 / 5

Workflow Clarity

The four steps (Triage → Mitigate → Communicate → Resolve) are clearly sequenced and include an explicit validation checkpoint ('Confirm error rates and latency return to baseline before declaring the incident contained') and time-boxed checklists. It misses 5 because there is no feedback loop for error recovery — no guidance when rollback fails or when no recent change can be identified.

4 / 5

Progressive Disclosure

The skill is a compact, self-contained SKILL.md (~55 lines) with no bundle files and no need for external references; sections are well-organized and template extraction to files (STATUS_UPDATE_TEMPLATE.md, POSTMORTEM_TEMPLATE.md) is explicitly signaled. Per the rubric's simple-skill exception, well-organized sections with no external-reference need warrant the top score.

5 / 5

Total

17

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete capabilities in third person, comprehensively covers the incident lifecycle, and provides an explicit 'Use when' clause rich in natural trigger terms with synonyms. The only weakness is mild overlap risk between performance-degradation triggers (error rate spikes, latency issues) and general debugging/observability skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions spanning the full incident lifecycle — 'alert acknowledgment, mitigation, status communication, and follow-up' plus 'rollback decisions, stakeholder updates, and structured post-incident review' — with comprehensive coverage and consistent third-person voice ('Manages', 'Guides'). It does not fit the 4 anchor because there are no meaningful coverage gaps across triage through postmortem.

5 / 5

Completeness

It explicitly answers both questions: a clear 'what' ('Manages production incident response workflows from initial triage through postmortem') and an explicit 'when' ('Use when the user mentions outages, service degradation, on-call pages...'). This matches the anchor for both what AND when with concrete trigger phrases.

5 / 5

Trigger Term Quality

The 'Use when' clause includes natural phrases users actually say: 'outages, service degradation, on-call pages, incident response, postmortems, error rate spikes, latency issues, or on-call escalation', with synonym variants (outages/service degradation, on-call pages/on-call escalation). Only trivial omissions like 'downtime' or 'site is down' keep it from being exhaustive, which still fits the comprehensive anchor.

5 / 5

Distinctiveness Conflict Risk

The incident-response niche is clear and mostly distinct, but the triggers 'error rate spikes, latency issues' could plausibly activate a general debugging or observability skill instead. It is above the 3 anchor (the framing is unmistakably incident-response) but does not fully meet the minimal-conflict bar of 5.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
fernandezbaptiste/skill-review-sandbox
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.