Expert OpenTelemetry guidance for collector configuration, pipeline design, and production telemetry instrumentation across Kubernetes, ECS, serverless, and standalone deployments. Use when configuring collectors, designing pipelines, instrumenting applications, implementing sampling, managing cardinality, securing telemetry, writing OTTL transformations, or setting up AI coding agent observability (Claude Code, Codex, Gemini CLI, GitHub Copilot).
97
95%
Does it follow best practices?
Impact
98%
1.10xAverage score across 21 eval scenarios
Passed
No findings from the security scan
Use this reference for Python SDK setup, framework instrumentation, async or streaming AI applications, and log-based events. For coding-agent products such as Claude Code or Gemini CLI, use ai-agents.md.
opentelemetry-distro, an OTLP exporter, and the matching
framework instrumentors; launch with opentelemetry-instrument. Run bootstrap
during image/environment construction, then lock the resolved dependencies.The upstream FastAPI/Starlette coexistence issue reports a risk of duplicate telemetry when native middleware and contrib instrumentation overlap. It does not establish native support in every released FastAPI/Starlette version. Verify the installed versions and one exported SERVER span per request before claiming this combination works. ASGI send/receive child spans are not duplicate SERVER spans.
For an explicit Python 1.44 example baseline:
python -m pip install opentelemetry-sdk==1.44.0 \
opentelemetry-exporter-otlp-proto-grpc==1.44.0
export OTEL_SERVICE_NAME=my-python-service
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317
export OTEL_EXPORTER_OTLP_PROTOCOL=grpcThe HTTP endpoint above is a local development example. Use TLS and the chosen exporter's authentication settings across trust boundaries. For HTTP/protobuf, install the HTTP exporter and use port 4318; do not send gRPC to an HTTP endpoint. Keep SDK/API/exporter versions aligned and let the instrumentor's dependency constraints determine its compatible beta version; do not give every package the SDK version number. Python 1.44 corresponds to the 0.65b0 release train.
The 1.44.0 release changes these application-facing behaviors:
LogRecord with its
event_name field, not just an attributes["event.name"] entry. Existing
backend attribute aliases are a separate compatibility concern.opentelemetry-configuration package (opentelemetry.configuration).
opentelemetry-sdk[file-configuration] remains an installation alias.OTEL_CONFIG_FILE, the environment-based initialization path is skipped;
OTEL_PYTHON_* extensions are bypassed. Put the intended settings in the file
and verify emitted resources, active instrumentors, and exporter destinations.
This is not a universal precedence rule for every exporter environment variable.ProcessResourceDetector no longer captures process.command_args or
process.command_line by default. Do not re-enable include_command_args=True
merely to restore a dashboard: arguments can contain credentials or user input.OTEL_LOGRECORD_ATTRIBUTE_COUNT_LIMIT and
OTEL_LOGRECORD_ATTRIBUTE_VALUE_LENGTH_LIMIT on the environment path. These
bound attributes, not arbitrary log bodies or sensitive-content exposure.Python GenAI development now lives in opentelemetry-python-genai.
Old contrib instrumentation-genai packages are deprecated and receive security
patches only. Read the destination package's README before replacing a package:
names, imports, configuration, and emitted attributes can change. Do not install
old and replacement instrumentors together.
Inventory checked September 10, 2026; these are beta, not stability guarantees:
| Library/framework | Destination package | Released version | Supported library range |
|---|---|---|---|
| OpenAI Python | opentelemetry-instrumentation-genai-openai | 1.1b0 | openai >=1.26.0,<4 |
| Anthropic Python | opentelemetry-instrumentation-genai-anthropic | 1.1b1 | anthropic >=0.51.0,<2 |
| Google GenAI | opentelemetry-instrumentation-google-genai | 1.1b1 | google-genai >=1.32.0,<3 |
| LangChain | opentelemetry-instrumentation-genai-langchain | 1.1b1 | langchain >=0.3.21,<2 |
| OpenAI Agents | opentelemetry-instrumentation-genai-openai-agents | 1.1b0 | openai-agents >=0.3.3,<1 |
Upstream lists Claude Agent SDK, CrewAI, DSPy, and LlamaIndex instrumentations as unreleased skeletons at this snapshot. A folder in the repository is not proof of an installable implementation. Verify package release and supported library range when selecting instrumentation. The ranges above do not imply validation of every SDK/framework combination by this skill.
Prefer an instrumentor that already covers your operations. Use the manual patterns below only for missing coverage; layering another inference span around an instrumented client will double-count operations and possibly token metrics.
GenAI conventions are Development. Pin the source revision used by your instrumentation/dashboard contract and preserve native fields during migration.
chat {model} (or the actual operation, such as
generate_content), with CLIENT kind for a remote model. Set provider identity
in gen_ai.provider.name; do not substitute the agent product name.invoke_agent {agent name} span.
Inference and subsequent tool execution are separate children of that invocation.execute_tool {tool name}, INTERNAL kind, and
gen_ai.operation.name=execute_tool plus gen_ai.tool.name. Use the registered
tool name, not arguments, paths, or request IDs. Do not rewrite vendor-native
spans solely to force this naming template.asyncio task creation. Preserve or explicitly
propagate context at detached tasks, threads, queues, and process boundaries.
Create tasks inside the intended parent scope and await owned work before ending it.The executable manual example uses small
provider-neutral adapters and never records response content. Real provider
adapters map usage into its input/output keys and expose an async iterator with
aclose(); cleanup must close the underlying provider stream.
from examples.python_telemetry import operation, streaming_inference, execute_tool
# tracer is obtained from the application's provider; chunks_factory is an
# adapter around the selected provider's streaming API, created inside the span.
async def run_agent(tracer, chunks_factory):
with operation(tracer, "invoke_agent assistant", attributes={
"gen_ai.operation.name": "invoke_agent", "gen_ai.agent.name": "assistant",
}):
async with streaming_inference(
tracer, chunks_factory, provider="example", model="example-model",
) as chunks:
async for chunk in chunks:
consume_in_application(chunk) # application code; no content telemetry
execute_tool(tracer, "lookup", lookup) # application callableThe example leaves intentional asyncio.CancelledError status UNSET and records
only an error class on other failures, re-raising the original exception. Classify
application timeouts separately from intentional cancellation. The outer context
manager closes the stream even when a consumer exits early.
from opentelemetry._logs import LogRecord
# Obtain this logger from the application's configured LoggerProvider.
# Run while the intended span is current; LogRecord captures its context.
def completed(logger):
logger.emit(LogRecord(event_name="agent.completed", attributes={"success": True}))Python 1.44 exposes these log APIs under _logs; verify imports when changing the
pin. The executable example includes complete trace/log providers, OTLP exporters,
and flushing. Production services should use batch processors, own providers once
per worker, and flush/shut down during graceful termination. Do not flush every
request or share an initialized exporter across a pre-fork worker boundary.
Run the isolated tests without model credentials or a Collector:
uv run --no-project --python 3.12 \
--with opentelemetry-sdk==1.44.0 \
--with opentelemetry-exporter-otlp-proto-grpc==1.44.0 \
--with opentelemetry-instrumentation-fastapi==0.65b0 \
--with fastapi==0.116.1 --with httpx==0.28.1 \
python -m unittest discover -s tests -p 'test_python_telemetry.py'Before rollout, inspect serialized OTLP at the Collector/backend: one server boundary per request, correct parentage, accurate span duration, log trace/span IDs, an event name field, and no sensitive content. In-memory tests alone do not prove the backend mapping or exporter configuration.
.claude-plugin
.codex-plugin
.cursor-plugin
.github
examples
references
tests