Expert OpenTelemetry guidance for collector configuration, pipeline design, and production telemetry instrumentation across Kubernetes, ECS, serverless, and standalone deployments. Use when configuring collectors, designing pipelines, instrumenting applications, implementing sampling, managing cardinality, securing telemetry, writing OTTL transformations, or setting up AI coding agent observability (Claude Code, Codex, Gemini CLI, GitHub Copilot).
97
95%
Does it follow best practices?
Impact
98%
1.10xAverage score across 21 eval scenarios
Passed
No findings from the security scan
otelcol validate --config <path> with the target distribution/version; render Helm values before validating the resulting config. For SDK changes, exercise a representative request and inspect emitted telemetry. Fix reported failures and repeat the affected check; if blocked, report the exact failure and remaining verification.memory_limiter first in every pipeline to prevent OOM crashes; size it below the container/host limit with runtime and buffer headroom. Prefer backpressure or controlled drops over crashing the Collector.user_id, request_id, session.id, or trace_id dimensions and explain time-series growth. Suggest privacy-safe traces/logs plus bounded metric labels; use the Rule of 100 in instrumentation.md as the review heuristic.traceID for tail sampling, tenant_id or cluster for tenant/shard routing); normalize non-string attributes first.load_balancing with routing_key: traceID and a Headless Service (clusterIP: None). For keep-errors-plus-10%, include error and probabilistic policies. Explicitly warn that tail_sampling is Beta in Collector 0.160; recheck stability for other releases.For coding-agent requests, load ai-agents.md and apply these signal-specific constraints:
CLAUDE_CODE_ENABLE_TELEMETRY=1, OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE=cumulative, and managed-settings persistence; traces are beta, while metrics and logs/events remain the broadly documented signals. Keep prompt/tool-content capture disabled unless PII controls are explicit, and avoid session.id as a metric dimension.gen_ai.* conventions, set gen_ai.operation.name: execute_tool, preserve gen_ai.tool.name, and use execute_tool {gen_ai.tool.name} for manually generated tool spans; keep the registered tool name in the attribute as well. These GenAI conventions are Development; preserve vendor-native spans.gen_ai.provider.name for the model/provider (for example, openai or gcp.gen_ai); use service.name or a natively emitted agent attribute for the coding-agent identity rather than writing the agent name into the provider field.Check these interactions, beyond configuration syntax:
tail_sampling, span_metrics, and service_graph need sticky routing above one replica.hostPort fits DaemonSet/node-local patterns, not scaled gateway Deployments.replicaCount, HPA minReplicas, PodDisruptionBudget, and rolling updates together.deltatocumulative / cumulativetodelta need source/backend/restart checks.file_storage needs local locking-safe storage; RWX, EFS, and NFS are unsafe defaults.Load only the references matching the request:
| Trigger keywords | Load | Key topics |
|---|---|---|
| New Collector pipeline, baseline config | production-baseline.md | Complete OTLP config, required substitutions, storage and listener setup |
| Choose deployment platform | setup-index.md | Platform decision matrix |
| Deploy to Kubernetes, Helm values.yaml | setup-kubernetes.md | DaemonSet, Gateway and sidecar manifests |
| Deploy to ECS, Fargate | setup-ecs.md | Task definitions, IAM and secrets |
| Deploy to Docker, Compose | setup-docker.md | Container networking, resources and volumes |
| Deploy to VM, EC2, systemd, Windows service | setup-vm.md | Service lifecycle and host setup |
| Kubernetes, Helm, values.yaml, audit, review, DaemonSet, Sidecar, Gateway, Scaling, Load Balancing | architecture.md | DaemonSet vs Gateway vs Sidecar, Target Allocator, HPA, rollout consistency |
| Pipeline, Receiver, Processor, Exporter, Queue, Batch, Memory, Extensions, existing config | collector.md | Processor ordering, memory_limiter, file_storage, config audit heuristics, temporality/state audits, stability levels |
| Python, FastAPI, Starlette, asyncio, Python GenAI, Python SDK events | python-instrumentation.md | Initialization ownership, duplicate instrumentation, SDK 1.44 migration, streaming, GenAI packages |
| SDK, Instrumentation, Spans, Attributes, Semantic Conventions, Cardinality | instrumentation.md | Auto vs manual, SemConv, cardinality Rule of 100 |
| Sampling, Cost, Volume, Head Sampling, Tail Sampling, Probabilistic | sampling.md | Head/tail sampling, sticky sessions, sampling math |
| Security, PII, GDPR, Redaction, TLS, Authentication, Credentials | security.md | PII redaction, mTLS, RBAC, extension exposure risks |
| Monitor the collector, Health, Alerts, Self-monitoring, Collector metrics | monitoring.md | otelcol_* metrics, dashboards, alert rules |
| Lambda, Azure Functions, GCP Functions, Serverless, FaaS, Mobile, Browser | platforms.md | FaaS patterns, Lambda extension layer, client-side apps |
| OTTL, Transform, Transformation, Modify, Filter attributes, Parse, Extract | ottl.md | OTTL syntax, context types, built-in functions, error handling |
| Connector, span_metrics, service_graph, signal_to_metrics, log-to-metric, span-to-metric, routing connector, failover connector | connectors.md | R.E.D. metrics, service graph, routing, failover, stickiness, generated-metric producer identity |
| Claude Code, Codex, Gemini CLI, Copilot, AI agent, coding agent, MCP | ai-agents.md | Agent OTel support matrix, unified collector config, GenAI SemConv |
| validate, dry-run, startup error, pipeline error, dropped data, queue full, recovery | validation.md | Config validation commands, live checks, symptom→cause→fix recovery guidance |
| playbook, production playbook, blog, developer observability, local OTel viewer, real world | playbooks.md | Production and developer patterns from OpenTelemetry and CNCF blogs |
| anti-pattern, common mistake, what to avoid, pitfall | anti-patterns.md | Full annotated anti-pattern catalogue: pipeline, metrics, Kubernetes, AI agents, OTTL |
.claude-plugin
.codex-plugin
.cursor-plugin
.github
examples
references
tests