CtrlK
BlogDocsLog inGet started
Tessl Logo

debug-inference

Troubleshoot failed or slow InferenceService deployments on OpenShift AI. Use when: - "My InferenceService won't start" - "Model deployment is stuck" - "Inference endpoint returns errors" - "Model is slow / high latency" - "GPU scheduling failed for my model" Progressive diagnosis: status conditions, events, pod logs, GPU health, and observability analysis. NOT for deploying models (use /model-deploy). NOT for creating runtimes (use /serving-runtime-config).

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong operational skill: concrete tool invocations with exact parameters, explicit OpenShift fallbacks for every rhoai call, human-in-the-loop checkpoints at each of six steps, and verification steps closing the loop. The main weaknesses are mild redundancy between the Prerequisites, Workflow, and Dependencies sections, and reference files that are pointer stubs rather than actual content, which slightly weakens the progressive-disclosure structure.

Suggestions

Move the inlined 'Issue 1: S3 Storage Access Denied' and 'Issue 2: NIM Authentication / GPU Incompatibility' sections into references/common-issues.md and keep only a one-line index in SKILL.md, mirroring how GPU/OOMKilled/RBAC issues are already delegated.

Trim the duplicated MCP tool inventory by letting the Prerequisites section (or a references file) be the single source of truth, and have the Workflow steps reference tools by name only.

Replace the pointer-stub reference files (e.g., common-issues.md containing '../../references/common-issues.md') with actual content or real symlinks so a reader following the reference lands on the material in one hop.

DimensionReasoningScore

Conciseness

The body is largely operational (tool calls with exact parameters, fallback paths, output templates) and avoids teaching concepts Claude already knows, fitting 'efficient; minor instances of over-explanation that could be trimmed'. Not a 5 because of minor redundancy: the full tool catalog in Prerequisites partially overlaps the Workflow and Dependencies sections, and two common issues (S3 access, NIM auth) are inlined even though common-issues.md exists as the designated home for them. Not a 3, since none of it is concept padding or filler — nearly every line is execution guidance.

4 / 5

Actionability

Guidance is fully executable: every step names the exact MCP tool with concrete parameters (apiVersion: serving.kserve.io/v1beta1, labelSelector: "serving.kserve.io/inferenceservice=[isvc-name]", container: "kserve-container"), specifies fallback tools when rhoai fails, includes a worked korrel8r query and sample PromQL, and ends with concrete verification steps. This matches 'fully executable; copy-paste ready commands; specific examples cover the common cases'.

5 / 5

Workflow Clarity

A clear 6-step progressive sequence with an explicit "WAIT for user confirmation" checkpoint after every step, quick-assessment prompts between steps, a recovery loop for unrecognized errors (live doc lookup), a structured diagnosis summary with a per-category status table, and post-fix verification steps. This matches the top anchor ('explicit validation steps; feedback loops for error recovery; checklists'); the skill is diagnostic/read-only and explicitly forbids auto-modification, so the destructive-operation cap does not apply.

5 / 5

Progressive Disclosure

The body is a well-structured overview with six clearly signaled one-level references (skill-conventions.md, openshift-fallback-templates.md, live-doc-lookup.md, common-issues.md, known-model-profiles.md, supported-runtimes.md), all of which exist in references/. It falls short of the top anchor because the reference files are themselves pointer stubs to other bundle locations (an extra indirection hop rather than content), and some detail that belongs in the references (two full common issues, the complete tool inventory) is inlined in SKILL.md. It is clearly above the score-3 anchor, since navigation is explicit and the split is mostly appropriate.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: third-person statement of capability, a comprehensive quoted 'Use when' trigger list covering start/stuck/error/slow/GPU symptoms, a one-line method summary, and explicit NOT-for boundaries that route adjacent intents to other skills. It fully satisfies the what/when/distinctness requirements with no padding.

DimensionReasoningScore

Specificity

The description names the domain ("failed or slow InferenceService deployments on OpenShift AI") and lists multiple concrete diagnostic actions: "status conditions, events, pod logs, GPU health, and observability analysis" — comprehensive coverage of the troubleshooting surface. It matches the anchor 'lists multiple specific concrete actions; comprehensive coverage' rather than the score-4 anchor, which requires minor coverage gaps; none are evident across failure and slowness symptoms.

5 / 5

Completeness

It explicitly answers both what ("Troubleshoot failed or slow InferenceService deployments... Progressive diagnosis: status conditions, events, pod logs, GPU health, and observability analysis") and when (a dedicated 'Use when:' list of concrete trigger phrases), matching the top anchor exactly. It is not the score-4 case, since the 'when' is fully explicit with quoted triggers rather than merely present.

5 / 5

Trigger Term Quality

The 'Use when' list quotes five natural user utterances with synonym coverage — "My InferenceService won't start", "Model deployment is stuck", "Inference endpoint returns errors", "Model is slow / high latency", "GPU scheduling failed for my model" — spanning start/stuck/slow/latency/error/GPU variations. This fits the 'comprehensive coverage of natural terms including synonyms' anchor; a user with any of these symptoms would naturally say one of these phrases.

5 / 5

Distinctiveness Conflict Risk

The niche is sharply delimited by the platform-specific trigger (InferenceService/OpenShift AI) and reinforced with explicit negative boundaries: "NOT for deploying models (use /model-deploy)" and "NOT for creating runtimes (use /serving-runtime-config)". Conflict risk with adjacent deployment/monitoring skills is minimal, matching the 'clear niche with distinct triggers' anchor.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
RHEcosystemAppEng/agentic-plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.