dd-trace-py LLMObs integration development guide. Use when creating, modifying, or debugging LLMObs integrations for LLM/AI libraries in the Python tracer. Covers BaseLLMIntegration, stream handling, message extraction, token counting, tool call parsing, and VCR-based testing patterns. Triggers: "llmobs", "LLMObs", "BaseLLMIntegration", "llmobs_set_tags", "_llmobs_set_tags", "BaseStreamHandler", "submit_to_llmobs", "integration.trace", "LLM span", "VCR", "cassette", "anthropic", "openai", "google_genai", "claude_agent_sdk", "generative-ai", "LLM integration", "llmobs_enabled".
LLMObs integrations enable Datadog LLM Observability for AI/LLM libraries. They extract model inputs, outputs, token usage, and tool calls from traced spans. This skill should be used in addition to the apm-integrations skill.
LLMObs integrations consist of two cooperating layers:
ddtrace/contrib/internal/{name}/patch.py) -- wraps library functions. Standard request/response LLM integrations construct LlmRequestEvent and use core.context_with_event() so the LLM tracing subscriber owns span lifecycle and LLMObs tag extraction.ddtrace/llmobs/_integrations/{name}.py) -- extends BaseLLMIntegration, implements _set_base_span_tags() and _llmobs_set_tags() to extract and set provider-specific messages, tools, metadata, and token metrics.Both layers must work together. The patch layer identifies the operation and passes request/response data through the event; the integration layer controls what data is extracted.
LlmRequestEvent with core.context_with_event() for new standard request/response LLM integrations. Anthropic is the canonical reference. This is the preferred pattern.integration.trace() and integration.llmobs_set_tags() directly, especially for child spans, agent/tool spans, or integrations not yet migrated. Google GenAI, OpenAI tool spans, and Claude Agent SDK are useful references.| Purpose | File |
|---|---|
| Base LLM integration class | ddtrace/llmobs/_integrations/base.py (BaseLLMIntegration) |
| Stream handler base classes | ddtrace/llmobs/_integrations/base_stream_handler.py (BaseStreamHandler, StreamHandler, AsyncStreamHandler) |
| Shared utilities | ddtrace/llmobs/_integrations/utils.py |
| LLMObs annotation helper | ddtrace/llmobs/_utils.py (_annotate_llmobs_span_data) |
| LLMObs constants | ddtrace/llmobs/_constants.py |
| LLMObs types | ddtrace/llmobs/types.py (Message, AudioPart, ToolCall, ToolResult, ToolDefinition) |
| Integration registry | ddtrace/llmobs/_integrations/__init__.py |
Always read 1-2 references before writing or modifying LLMObs code.
| Provider | Patch File | LLMObs Integration | LLMObs Tests |
|---|---|---|---|
| Anthropic (canonical) | ddtrace/contrib/internal/anthropic/patch.py | ddtrace/llmobs/_integrations/anthropic.py | tests/contrib/anthropic/test_anthropic_llmobs.py |
| Claude Agent SDK (latest, agent pattern) | ddtrace/contrib/internal/claude_agent_sdk/patch.py | ddtrace/llmobs/_integrations/claude_agent_sdk.py | tests/contrib/claude_agent_sdk/test_claude_agent_sdk_llmobs.py |
| OpenAI | ddtrace/contrib/internal/openai/patch.py | ddtrace/llmobs/_integrations/openai.py | tests/contrib/openai/test_openai_llmobs.py |
| Google GenAI | ddtrace/contrib/internal/google_genai/patch.py | ddtrace/llmobs/_integrations/google_genai.py | tests/contrib/google_genai/test_google_genai_llmobs.py |
Use Anthropic as the canonical reference for standard LLM integrations. Use Claude Agent SDK for agent-pattern integrations (agent spans, tool child spans, thinking blocks).
Subclass BaseLLMIntegration and implement:
_set_base_span_tags(span, **kwargs)Set provider-specific APM tags on the span (e.g., {name}.request.model).
_llmobs_set_tags(span, args, kwargs, response, operation)Extract and annotate all LLMObs fields on the span:
| Field | Description |
|---|---|
kind | "llm" for LLM calls, "agent" for agent calls, "tool" for tool calls |
model_name | Model identifier (e.g., "claude-3-sonnet-20240229") |
model_provider | Provider name (e.g., "anthropic", "openai") |
input_messages | List of Message objects from request |
output_messages | List of Message objects from response |
metadata | Dict of sanitized request parameters (temperature, top_p, etc.), plus response-derived scalars such as finish_reason (see Response Metadata below) |
metrics | Token usage dict with INPUT_TOKENS_METRIC_KEY, OUTPUT_TOKENS_METRIC_KEY, TOTAL_TOKENS_METRIC_KEY |
tool_definitions | List of ToolDefinition objects if tools are passed |
Fields are usually set via _annotate_llmobs_span_data(...), not raw span._set_ctx_items(...).
metadata is not request-only. Scalars the provider reports on the response belong there too, so
long as they are not already covered by metrics (token counts) or output_messages.
The stop reason is the established case. Record it under the key finish_reason for every provider,
so one facet answers "why did generation stop" regardless of integration, and keep each provider's
own vocabulary as the value (stop/length/content_filter/tool_calls for the OpenAI family,
end_turn/max_tokens/refusal/tool_use for Anthropic). Read it from:
| Provider | Source |
|---|---|
| openai/litellm chat + legacy completions | choice.finish_reason (per choice) |
| openai responses API | incomplete_details.reason |
| anthropic | response.stop_reason, falling back to response.finish_reason for the streamed dict the aggregator rebuilds |
Rules:
parameters.update(...)), so response data never clobbers request params.finish_reason out
entirely rather than writing None or "".n > 1), comma-join
the per-choice reasons in choice order ("stop,length") instead of emitting a list, so the key
never changes type. _openai_finish_reason_metadata() in _integrations/utils.py does this.span.error. A response can exist on an errored span (see the AI
Guard case in _integrations/anthropic.py); gate on response is not None.Note two already-shipped integrations predate this key: bedrock and the claude-agent-sdk both write
metadata["stop_reason"]. Renaming those is a breaking change and has not been done — follow
finish_reason for new work.
submit_to_llmobs=True must be set on LlmRequestEvent for event-based request spans or passed to integration.trace() for direct LLMObs spansctx.dispatch_ended_event() must run on success and error paths for event-based patch wrappersBaseStreamHandler/AsyncStreamHandler -- never consume streams directlyspan.set_exc_info(), span.finish(), or integration.llmobs_set_tags() directly; the tracing subscriber handles that when the event endsintegration.llmobs_set_tags() and span lifecycle handling aligned with the closest current referencemodule._datadog_integration = MyLibIntegration(integration_config=config.mylib)from ddtrace.llmobs.types import AudioPart, ImagePart, Message, ToolCall, ToolResult, ToolDefinition
# Input/output messages
Message(content="text", role="user")
Message(content="response", role="assistant", tool_calls=[...])
# Audio attachments in multimodal messages
AudioPart(mime_type="audio/wav", content="<base64-audio>")
Message(content="", role="user", audio_parts=[...])
# Image attachments in multimodal messages
ImagePart(mime_type="image/png", content="<base64-image>")
Message(content="", role="user", image_parts=[...])
# Tool calls (in output messages)
ToolCall(name="get_weather", arguments={"city": "NYC"}, tool_id="toolu_123", type="tool")
# Tool results (in input messages)
ToolResult(result="72F sunny", tool_id="toolu_123", type="tool_result")
# Tool definitions (from request parameters)
ToolDefinition(name="get_weather", description="...", schema={...})submit_to_llmobs=True, ctx.dispatch_ended_event(), and llmobs_enabledBaseStreamHandler subclass, check finalize_stream() dispatches the ended event or finishes direct-trace spans according to the reference patternSee Failure Modes for detailed debugging guide.
57aff59
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.