CtrlK
BlogDocsLog inGet started
Tessl Logo

k8s-troubleshooter

Systematic Kubernetes troubleshooting and incident response. Use this skill whenever the user mentions Kubernetes, K8s, kubectl, pods, containers, or clusters. Triggers include diagnosing CrashLoopBackOff, ImagePullBackOff, OOMKilled, or Pending pods, responding to production incidents, troubleshooting node NotReady or DiskPressure, debugging service connectivity or networking, investigating PVC or storage failures, analyzing performance degradation, checking cluster health, troubleshooting Helm releases, and conducting post-incident reviews.

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-structured, highly actionable troubleshooting workflow with excellent progressive disclosure to four reference files and a script. Its main weakness is repetition of the same diagnostic commands across sections, which inflates token use without adding value.

Suggestions

Consolidate the duplicated kubectl pod/node diagnostic commands into a single Quick Reference section and reference it elsewhere rather than repeating the full command blocks in the Core Workflow and Diagnostic Scripts sections.

Remove or tighten narrative restatement lines like 'This provides an overview of:' and 'This reveals:' since Claude can infer what the commands output.

Add an explicit validate-then-proceed feedback loop in Step 5/6 for production remediation (e.g., re-run the namespace health check, confirm the pod is Running, and roll back if the fix fails verification) to strengthen the workflow for destructive/batch operations.

DimensionReasoningScore

Conciseness

The body is mostly efficient with concrete commands, but it repeats the same kubectl diagnostics (describe pod, logs, get events) and check_namespace.py usage across multiple sections, and includes minor over-explanatory lines like "This provides an overview of:" and "This reveals:" that could be trimmed.

3 / 5

Actionability

It provides copy-paste-ready kubectl commands and a documented Python script with flags (--json, --events 20), covering the common pod/node/service/storage cases fully and executably.

5 / 5

Workflow Clarity

A clear six-step sequence runs from Gather Context through Verify & Monitor with an explicit verification step, but remediation in production lacks a structured validate-fix-then-retry feedback loop beyond general monitoring.

4 / 5

Progressive Disclosure

The body is a concise overview that clearly signals one-level-deep references (common_issues.md, incident_response.md, performance_troubleshooting.md, helm_troubleshooting.md) and the check_namespace.py script, each with a 'Read this when...' navigation cue; all referenced files exist.

5 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: it clearly states the skill's purpose, lists comprehensive concrete capabilities, and provides explicit, natural-language triggers with synonyms. It fully answers both 'what' and 'when' with minimal conflict risk.

DimensionReasoningScore

Specificity

The description enumerates many concrete actions ("diagnosing CrashLoopBackOff, ImagePullBackOff, OOMKilled", "troubleshooting node NotReady or DiskPressure", "investigating PVC or storage failures", "troubleshooting Helm releases"), giving comprehensive coverage of the skill's capabilities.

5 / 5

Completeness

It explicitly answers both what ("Systematic Kubernetes troubleshooting and incident response") and when ("Use this skill whenever the user mentions Kubernetes, K8s, kubectl, pods, containers, or clusters"), with concrete trigger phrases.

5 / 5

Trigger Term Quality

It includes the natural terms a user would say ("Kubernetes, K8s, kubectl, pods, containers, or clusters") plus specific error states, with synonyms like K8s, matching the comprehensive-synonym anchor.

5 / 5

Distinctiveness Conflict Risk

It carves a clear Kubernetes-troubleshooting niche with distinct triggers unlikely to fire for unrelated skills, minimizing conflict risk.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
ahmedasmar/devops-claude-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.