CtrlK
BlogDocsLog inGet started
Tessl Logo

otel-genai-instrumentation

Guides instrumentation of GenAI/LLM applications with OpenTelemetry for Honeycomb, including content capture and agent failure detection. Trigger phrases: "instrument my GenAI app", "add tracing to LLM calls", "trace AI agent", "instrument OpenAI", "instrument Anthropic", "GenAI observability", "trace tool calling", "LLM token usage", "instrument embeddings", "trace MCP", "GenAI metrics", "instrument LangChain", "add GenAI spans", "capture prompts", "capture LLM responses", "enable GenAI content capture", "streaming tracing", "trace streaming responses", or any request about instrumenting GenAI/LLM applications.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured skill with strong progressive disclosure and a clearly sequenced critical-requirements workflow. Its main weakness is conciseness: the duplicated content-capture prompt and repeated attribute/impact lists inflate the body without adding information.

Suggestions

Remove the verbatim duplicate of the content-capture question: state it once in the Content Capture section and reference it from the Critical Requirements checklist instead of repeating the full block.

Consolidate the repeated required-attributes and 'impact if missing' material (Steps 3, the Required Attributes section, and the Attribute Completeness impact list) into a single authoritative location and cross-reference it.

Add a short in-flow validation step to the Critical Requirements workflow (e.g., confirm a GenAI span with gen_ai.operation.name appears in Honeycomb) rather than deferring all verification to the query-patterns skill.

DimensionReasoningScore

Conciseness

The body is mostly efficient and avoids explaining basic concepts, but it repeats material verbatim — the full content-capture question appears twice (Critical Requirements Step 1 and Content Capture Step 1) and the required-attributes/impact-if-missing lists recur several times — which is genuine tighten-able padding rather than lean token use.

3 / 5

Actionability

It provides copy-paste-ready code in Python/Node.js/Go for span flushing, exact environment-variable exports, package tables with minimum version pins, and precise span-name and attribute patterns, covering the common cases fully.

5 / 5

Workflow Clarity

The 'Critical Requirements (Non-Negotiable)' section gives an explicit ordered checklist (ask about content capture FIRST, enable conventions, set required attributes, implement force_flush) with impact-if-missing feedback, but final verification of GenAI spans flowing is deferred to other skills rather than an in-flow validation checkpoint.

4 / 5

Progressive Disclosure

The body is a clear overview that delegates detail to seven well-signaled, one-level-deep reference files (all of which exist in ./references/), with an 'Additional Resources' index summarizing each — matching the clear-overview, easy-navigation anchor.

5 / 5

Total

17

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-structured description that gives a clear capability statement and a rich set of natural trigger phrases covering both what and when. Minor tightening of the capability list and a touch more disambiguation from the base OTEL skill would push it to the top anchor.

DimensionReasoningScore

Specificity

The description names the domain and several concrete capabilities ("content capture", "agent failure detection") plus a wide set of specific action verbs in the trigger phrases (instrument OpenAI/Anthropic/LangChain, trace MCP, instrument embeddings). It stops short of a fully comprehensive capability list, sitting at 'several specific actions; minor gaps'.

4 / 5

Completeness

It explicitly answers both 'what' (guides instrumentation of GenAI/LLM applications with OpenTelemetry for Honeycomb, including content capture and agent failure detection) and 'when' via concrete trigger phrases plus a catch-all ('or any request about instrumenting GenAI/LLM applications'), matching the anchor for explicit what-and-when with concrete triggers.

5 / 5

Trigger Term Quality

It provides an extensive list of natural phrases a user would actually say ("instrument my GenAI app", "add tracing to LLM calls", "GenAI observability", "LLM token usage") with synonyms and provider names, matching the comprehensive coverage anchor.

5 / 5

Distinctiveness Conflict Risk

It carves a clear GenAI/LLM-tracing niche with distinct triggers, but provider-agnostic phrases like "instrument OpenAI" and "add tracing to LLM calls" carry minor overlap risk with a closely related base otel-instrumentation skill, keeping it just below the minimal-conflict anchor.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (666 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
honeycombio/agent-skill
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.