CtrlK
BlogDocsLog inGet started
Tessl Logo

deploy-otel

Deploy the OpenTelemetry observability stack (Prometheus, Grafana, OTEL Collector) to a Kind cluster for testing toolhive telemetry. Use when you need to set up monitoring, metrics collection, or observability infrastructure.

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured runbook: every step is executable copy-paste bash with idempotency, verification, troubleshooting, and cleanup sections. The only meaningful gap is that final verification lists pods without a status check or error-recovery loop, and there is minor token overhead from ceremonial echo statements.

DimensionReasoningScore

Conciseness

The body is dominated by lean, executable bash blocks with no explanations of concepts Claude already knows (no "what is Prometheus" padding). Minor instances of trimming opportunity remain, such as ceremonial echo lines ("echo \"Checking prerequisites...\"", "echo \"All prerequisites met.\"") that add tokens without adding guidance, so it does not quite reach the every-token-earns-its-place anchor.

4 / 5

Actionability

Every step is copy-paste-ready, fully executable bash: prerequisite checks with explicit failure exits, idempotent cluster creation, helm commands with concrete values files ("-f examples/otel/prometheus-stack-values.yaml"), timeouts, port-forward instructions with credentials, and cleanup commands. This matches the anchor for fully executable guidance covering the common cases.

5 / 5

Workflow Clarity

Steps 1–8 are clearly sequenced, and validation checkpoints exist (prerequisite checks, helm --wait flags, step 7 "kubectl get pods -n monitoring" verification). However, verification stops at listing pods — there is no check that pods reach Running state or a validate-fix-retry feedback loop if they do not, which the 5 anchor requires; the gap is minor, so 4 fits.

4 / 5

Progressive Disclosure

The single-file body is well-organized into purposeful sections (Steps, Troubleshooting, What This Deploys, Cleanup) with no buried references and no nested-reference chains; chart documentation lives at one-level-deep external URLs. It is not under 50 lines, and sections like the Troubleshooting prose are slightly long for an overview file, which keeps it below the 5 anchor.

4 / 5

Total

17

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly covers both what it does and when to use it, with concrete component names and natural trigger phrases. Its only notable flaw is the second-person phrasing in the trigger clause and a single-action framing that keeps specificity from scoring higher.

Suggestions

Rewrite the trigger clause in third person to avoid the second-person penalty, e.g., "Use when setting up monitoring, metrics collection, or observability infrastructure" instead of "Use when you need to...".

Add a few more natural trigger synonyms such as "tracing", "dashboards", or "deploy Prometheus/Grafana" to broaden keyword coverage for users phrasing the need differently.

DimensionReasoningScore

Specificity

"Deploy the OpenTelemetry observability stack (Prometheus, Grafana, OTEL Collector) to a Kind cluster" names a concrete action with specific components, which would merit a 4, but the trigger clause "Use when you need to set up monitoring" uses second-person voice, which the guidelines penalize by 1. It also describes essentially one action (deploy) rather than multiple concrete actions.

3 / 5

Completeness

Both questions are explicitly answered: what — "Deploy the OpenTelemetry observability stack (Prometheus, Grafana, OTEL Collector) to a Kind cluster for testing toolhive telemetry" — and when — "Use when you need to set up monitoring, metrics collection, or observability infrastructure". This matches the anchor for clear and explicit what AND when with concrete trigger phrases.

5 / 5

Trigger Term Quality

Natural terms like "monitoring", "metrics collection", and "observability infrastructure" are phrases users would actually say, alongside specific stack names (OpenTelemetry, Prometheus, Grafana). A few natural variations are missing (e.g., "tracing", "dashboards", "telemetry setup"), so it does not reach the comprehensive synonym coverage of a 5.

4 / 5

Distinctiveness Conflict Risk

The description occupies a clear niche — deploying an OTEL observability stack to a Kind cluster specifically for toolhive telemetry testing — with triggers (monitoring/metrics/observability setup) that are unlikely to fire for unrelated skills. Minimal conflict risk.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
stacklok/toolhive
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.