CtrlK
BlogDocsLog inGet started
Tessl Logo

prometheus-cardinality-troubleshooter

Diagnostic guide for active Prometheus cardinality problems — slow queries, OOMing Prometheus, high Grafana Cloud Active Series or DPM bills, "too many samples" ingest errors, series churn, or rapid memory growth. Walks through tsdb status endpoints, per-metric and per-label drill-downs, common-culprit galleries, and remediation paths. Use when the user is *currently experiencing* a cardinality fire. For preventing cardinality issues at the source, route to prometheus-label-strategy. For post-ingest aggregation, route to adaptive-metrics. For DPM-specific analysis, route to dpm-finder.

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality operational triage guide: copy-paste-ready PromQL/YAML/Alloy, a clear sequenced diagnostic workflow with routing tables and a decision tree, and well-structured navigation. The only soft spots are mild repetition of the core safety warning and the absence of an explicit post-remediation verification loop.

Suggestions

Add a short "Verify after remediation" feedback loop: after dropping a metric or applying Adaptive Metrics, re-check prometheus_tsdb_head_series / DPM and confirm rate()/increase() still return sane values, so the destructive-action workflow has an explicit validation checkpoint.

Consolidate the "One Rule" warning into its section and reference it once from the Common-Culprit Gallery and Emergency Drop Patterns instead of restating it verbatim, to recover the repeated tokens.

DimensionReasoningScore

Conciseness

The body is dense and operational — PromQL, curl|jq, YAML, and specific thresholds (">10K unique values", "14× multiplier", ">a few % per day") — and assumes Claude knows PromQL rather than re-explaining it, fitting the score-4 "efficient; minor instances that could be trimmed" anchor; it is not score 5 because the "One Rule" warning is deliberately restated across several sections, adding repeat tokens.

4 / 5

Actionability

Fully executable, copy-paste-ready guidance throughout — tsdb status curl, per-metric drill-down PromQL, kube-state-metrics flags, metric_relabel_configs YAML, and an Alloy component block under "Emergency Drop Patterns (copy-paste ready)" — covering the common cases per the score-5 anchor; no pseudocode gaps that would drop it to 4.

5 / 5

Workflow Clarity

Clear Step 1→5 sequence, a symptom→cause routing table, and a remediation decision tree, with an "Always test in staging first" checkpoint and conclusive go/no-go signals ("vertical step aligned with a deploy is conclusive"), matching score 4; not 5 because there is no explicit verify-after-drop feedback loop (confirm rate()/DPM still sane after a remediation), though validation is present enough that the destructive-operation cap at 3 does not apply.

4 / 5

Progressive Disclosure

Well-organized sections with clear internal anchor navigation (#step-1-active-series-triage, etc.) and a dedicated "When to Hand Off" lane; no bundle files exist so all content is inline, which fits score 4 ("good structure; minor organization gaps") rather than 5 (which expects one-level-deep external references) or 3 (structure is genuinely strong, not merely "could be better organized").

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, tightly-scoped description: it states concrete diagnostic actions, surfaces natural trigger phrases an SRE would actually say, explicitly answers both what and when, and carves clear boundaries against sibling skills. No first/second-person voice or vague fluff is present.

DimensionReasoningScore

Specificity

"Walks through tsdb status endpoints, per-metric and per-label drill-downs, common-culprit galleries, and remediation paths" lists multiple concrete diagnostic actions with comprehensive coverage, matching the score-5 anchor rather than the score-4 anchor which expects minor gaps.

5 / 5

Completeness

Explicitly answers both "what" (diagnostic walkthrough plus remediation paths) and "when" ("Use when the user is *currently experiencing* a cardinality fire") with concrete trigger phrases, matching the score-5 anchor; the "when" is fully explicit, not merely implied as in the score-4 anchor.

5 / 5

Trigger Term Quality

Natural SRE phrasing is comprehensive — "slow queries", "OOMing Prometheus", "high Grafana Cloud Active Series or DPM bills", "too many samples", "series churn", "rapid memory growth" — including the actual ingest error string users would quote; this exceeds the score-4 anchor ("a few natural terms missing").

5 / 5

Distinctiveness Conflict Risk

Clear niche (active cardinality fire, "diagnosis under pressure") with explicit boundary routing to sibling skills (prometheus-label-strategy for prevention, adaptive-metrics for post-ingest aggregation, dpm-finder for DPM), giving minimal conflict risk per the score-5 anchor rather than the "minor overlap" of score 4.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
grafana/skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.