CtrlK
BlogDocsLog inGet started
Tessl Logo

observability-rca

Use this skill when performing root cause analysis on incidents detected by Elastic Observability. Activate when the user reports a production issue, outage, degraded performance, or asks to investigate alerts.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, highly actionable RCA framework built around executable ES|QL queries and a clear numbered investigation sequence, with no wasted tokens and well-organized self-contained structure. The one gap is the absence of an explicit validation step confirming the root-cause hypothesis before moving to documentation.

Suggestions

Add an explicit hypothesis-validation checkpoint between Correlation Analysis and Resolution Documentation (e.g., 'Confirm the candidate root cause against at least two independent signals before documenting it').

Optionally note how to handle queries that return no results or ambiguous signals, giving a light feedback loop for dead-end investigation paths.

DimensionReasoningScore

Conciseness

The body is lean: numbered section headers plus executable ES|QL commands and a compact table, with only brief contextual intros that earn their place and no explaining of concepts Claude already knows.

5 / 5

Actionability

Provides copy-paste-ready ES|QL queries covering scope, timeline, correlation (traces/CPU/network), and deploy-change detection; the only placeholders (<affected-service>, <service>) are explicitly justified by incident-dependent values.

5 / 5

Workflow Clarity

A clear, well-numbered 5-step investigation sequence (assess scope → timeline → correlation → root causes → documentation) with most structure present, but it lacks an explicit hypothesis-verification checkpoint to confirm a candidate root cause before documenting it.

4 / 5

Progressive Disclosure

A self-contained, single-purpose skill with no need for external references, organized into five clearly numbered sections with easy navigation, matching the simple-skill exception for well-organized content.

5 / 5

Total

19

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-targeted description that clearly answers what the skill does and when to activate it, with natural trigger phrasing tied to a specific product niche. The main gap is capability specificity, which names only one action rather than enumerating the investigation activities the skill supports.

Suggestions

Expand specificity by listing the concrete RCA activities the skill performs (e.g., timeline reconstruction, cross-service correlation, infrastructure metric analysis) rather than a single 'root cause analysis' action.

Broaden trigger terms with common operational synonyms users say during incidents (e.g., '5xx errors', 'latency spike', 'error rate', 'on-call investigation', 'pager').

DimensionReasoningScore

Specificity

Names the domain and one concrete action ("performing root cause analysis on incidents"), but does not enumerate several specific capabilities, so it sits at the 'domain + 1-2 actions' anchor rather than a multi-action one.

3 / 5

Completeness

Explicitly states both what it does (RCA on Elastic Observability incidents) and when to use it ("Activate when the user reports a production issue, outage, degraded performance, or asks to investigate alerts") with concrete trigger phrases.

5 / 5

Trigger Term Quality

Includes natural trigger phrases a user would say ("production issue", "outage", "degraded performance", "investigate alerts") with good coverage, but misses common synonyms like 5xx, latency, error rate, or on-call/pager.

4 / 5

Distinctiveness Conflict Risk

Tied to a specific product (Elastic Observability) with distinct incident-oriented triggers, giving it a clear niche with minimal overlap risk against other skills.

5 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
elastic/elastic-ramen
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.