CtrlK
BlogDocsLog inGet started
Tessl Logo

model-deploy

Deploy AI/ML models on OpenShift AI using KServe with vLLM, NVIDIA NIM, or Caikit+TGIS runtimes. Use when: - "Deploy Llama 3 on my cluster" - "Set up a vLLM inference endpoint" - "Deploy a model with NIM" - "Create an InferenceService for Granite" - "I need to serve a model on OpenShift AI" Handles runtime selection, GPU validation, InferenceService CR creation, and rollout monitoring. NOT for NIM platform setup (use /nim-setup first). NOT for custom runtime creation (use /serving-runtime-config).

71

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced operational workflow with strong tool-level specificity and human-in-the-loop gating, undermined by content duplication (taint handling repeated three times, overlapping Step 1/2 questions) and by progressive-disclosure deferrals pointing at empty reference files. Tightening the body and actually populating the referenced files would lift both weak dimensions.

Suggestions

Populate the empty reference files (supported-runtimes.md, live-doc-lookup.md, openshift-fallback-templates.md, common-issues.md, skill-conventions.md) or move the inlined Common Issues and toleration patterns into them — Steps 3-4 currently direct the agent to read files that contain no content.

Merge the duplicate GPU-taint sections (Step 9 'GPU Toleration Handling', 'Issue 4', and the unnumbered 'Issue: Pod Stuck Pending Due to GPU Node Taints') into a single authoritative location, keeping the others as one-line pointers.

Deduplicate the Step 1/Step 2 question lists — 'Model source' and 'Deployment mode' are asked in both steps; ask each once and reference it in the configuration table.

DimensionReasoningScore

Conciseness

The body is mostly operational rather than explanatory, but contains real duplication: Step 1 and Step 2 both ask the user for "Model source" and "Deployment mode", and GPU-taint toleration content appears three times (Step 9 "GPU Toleration Handling", "Issue 4: GPU Node Taints Prevent Scheduling", and the near-duplicate unnumbered "Issue: Pod Stuck Pending Due to GPU Node Taints"). This is more than the 'minor instances that could be trimmed' of the 4 anchor, but not the pervasive padding of the 2 anchor — a solid 3.

3 / 5

Actionability

Fully executable guidance: exact MCP tool names with REQUIRED/OPTIONAL parameter lists and defaults, concrete apiVersion/kind/labelSelector values, a complete toleration YAML snippet, NIM fallback fields (NGC_API_KEY secretKeyRef, pinned image tag, NIM_MAX_MODEL_LEN), and copy-paste curl test commands covering both OpenAI-compatible and KServe v2 endpoints. The only placeholder content ([model-name], {"inputs":[...]}) is inherently user-specific, so this sits at the 5 anchor rather than the 4.

5 / 5

Workflow Clarity

An 11-step workflow with explicit validation checkpoints (pre-flight environment validation before configuration effort, explicit "WAIT for user confirmation" gates in Steps 1, 2, 3, 8, and on failure), a pre-flight checklist reference, rollout monitoring with polling intervals and timeout, and error-recovery option menus (pod logs, events, /debug-inference, retry). The destructive-operation cap does not apply: deletion is explicitly gated ("NEVER auto-delete failed deployments"), so this is a clear 5.

5 / 5

Progressive Disclosure

The reference structure is well signaled and one level deep, but 5 of the 8 referenced bundle files are empty (supported-runtimes.md, live-doc-lookup.md, openshift-fallback-templates.md, common-issues.md, skill-conventions.md are 0 bytes) while workflow steps 3-4 instruct reading them for essential content, and ~90 lines of Common Issues are inlined that the empty common-issues.md was meant to hold. Referenced anchors (#inferenceservice-nim, #deploy-model-missing-gpu-tolerations) resolve to nothing. The deferral pattern is right but the split is not realized — this is more than the 'minor organization gaps' of the 4 anchor.

3 / 5

Total

16

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete capability list with named technologies, five natural trigger phrases with synonym coverage, explicit what/when, and boundary clauses that route adjacent intents to sibling skills. No changes needed.

DimensionReasoningScore

Specificity

"Handles runtime selection, GPU validation, InferenceService CR creation, and rollout monitoring" lists multiple specific concrete actions spanning the full deployment lifecycle, and the opening line names the exact stack (KServe, vLLM, NVIDIA NIM, Caikit+TGIS). This matches the 5 anchor (comprehensive coverage); the 4 anchor would require minor gaps in action coverage, which are not present.

5 / 5

Completeness

It explicitly answers both what ("Deploy AI/ML models on OpenShift AI using KServe... Handles runtime selection, GPU validation, InferenceService CR creation, and rollout monitoring") and when (a dedicated "Use when:" list with concrete trigger phrases), plus boundary exclusions. The "Use when" clause is present, so no cap applies; this is a clean 5, not a 4, because the 'when' is fully explicit rather than improvable.

5 / 5

Trigger Term Quality

The "Use when" block provides five natural user utterances with synonym coverage ("Deploy Llama 3 on my cluster", "Set up a vLLM inference endpoint", "Deploy a model with NIM", "Create an InferenceService for Granite", "I need to serve a model on OpenShift AI") — deploy/set up/create/serve variants across all three runtimes, matching the 5 anchor's comprehensive synonym coverage. Not a 4: no common natural phrasing for this domain is conspicuously missing.

5 / 5

Distinctiveness Conflict Risk

A clear niche (KServe model serving on OpenShift AI) with distinct triggers and explicit "NOT for NIM platform setup (use /nim-setup first)" / "NOT for custom runtime creation (use /serving-runtime-config)" disambiguation, giving minimal conflict risk. Third-person voice throughout ("Deploy", "Handles"), so no specificity penalty applies.

5 / 5

Total

20

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 deeper-than-1-level

Warning

referenced_paths_exist

Referenced path issues: 1 deeper-than-1-level

Warning

Total

13

/

16

Passed

Repository
RHEcosystemAppEng/agentic-plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.