Trace CrewAI crews and flows with Maple: OpenInference CrewAI instrumentor plus the model-SDK instrumentor, GenAI dual-write, one Maple Agent Session per conversation with transcript, model and tool calls, tokens, failed tools and one lane per agent. Triggers on 'trace my crewai agent', 'add Maple to crewai', 'agent sessions for crewai', 'OpenTelemetry for crewai'.
76
95%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Goal: every conversation the app runs through CrewAI shows up in Maple Agent Sessions as ONE session, with the transcript, each model call (model, tokens), each tool call (name, result, failure) and one lane per agent role.
CrewAI exports nothing to your backend. All spans come from OpenInference, and the defaults are wrong for Maple in three ways this skill fixes: the CrewAI instrumentor records no model calls (a second, SDK-level instrumentor is required), no session id, and no agent-name attribute.
pyproject.toml / requirements*.txt / uv.lock / poetry.lock. Need crewai>=1.15, Python 3.10-3.13. If older, upgrade CrewAI first (1.x changed provider routing).LLM(model=...) / Agent(llm=...) string and map it to the SDK CrewAI calls. Install one instrumentor per SDK actually used:| Model string | SDK | Instrumentor package / class |
|---|---|---|
openai/…, openrouter/…, deepseek/…, ollama/…, hosted_vllm/…, cerebras/…, dashscope/…, custom_openai=True, or any bare name not matched below | openai | openinference-instrumentation-openai / OpenAIInstrumentor |
anthropic/…, claude/…, bare claude-… | anthropic | openinference-instrumentation-anthropic / AnthropicInstrumentor |
gemini/…, google/…, bare gemini-… | google-genai | openinference-instrumentation-google-genai / GoogleGenAIInstrumentor |
bedrock/…, aws/…, bare anthropic.claude-… | boto3 | openinference-instrumentation-bedrock / BedrockInstrumentor |
any other prefix (LiteLLM fallback, needs crewai[litellm]) | litellm | openinference-instrumentation-litellm / LiteLLMInstrumentor |
An agent with no llm= uses env MODEL / MODEL_NAME / OPENAI_MODEL_NAME, else gpt-4.1-mini (the openai row); resolve that string with the same table. azure/… uses azure-ai-inference, which has no OpenInference instrumentor: tell the user model calls won't be recorded.
kickoff call site (crew.kickoff, kickoff_async, akickoff, flow.kickoff, flow.handle_turn, flow.resume, Agent.kickoff) and how conversations are identified (chat id, thread id, session row, flow state.id).TracerProvider, trace.set_tracer_provider, opentelemetry-instrument, logfire.configure, phoenix.otel.register, langfuse, sentry_sdk.init, CrewAIInstrumentor, litellm.callbacks = ["otel"]. If a provider exists, REUSE it: add Maple's exporter and the processor below to it and pass it to instrument(). Never create a second provider. If an instrumentor's instrument() already runs, change that call instead of adding another.OTEL_SDK_DISABLED. If set to true, remove it (it kills the whole SDK) and replace with CREWAI_DISABLE_TELEMETRY=true.https://ingest.maple.dev. EU: https://ingest.eu.maple.dev.Authorization=Bearer <key>. Protocol: http/protobuf.MAPLE_TEST (ingest accepts and discards it) and tell the user to replace it with their key from Settings → Ingestion.maple_sk_ key in browser code..env, settings module, secret manager) if it has one. Otherwise inlining the key is acceptable: ingest keys are write-only..env (load_dotenv()): call it at the top of tracing.py, before the provider is built. Otherwise the exporter silently targets localhost:4318 with no key.tracing.py, after any load_dotenv(), when OTEL_EXPORTER_OTLP_HEADERS is unset, log one warning (logging.getLogger(__name__).warning("OTEL_EXPORTER_OTLP_HEADERS (Maple ingest key) is not set; Maple telemetry export is disabled")) and skip the provider and exporter setup. Never raise or exit over the key, and never send a header without one (opaque 401).pip install "crewai>=1.15" "openinference-instrumentation-crewai>=1.1.18" \
"openinference-instrumentation-openai>=0.1.61" \
"opentelemetry-sdk>=1.45" "opentelemetry-exporter-otlp-proto-http>=1.45"Use the repo's package manager (uv add, poetry add...). Swap/add the model-SDK instrumentor per the Step 0 table.
Environment (in the repo's env mechanism):
OTEL_SERVICE_NAME=<service name, e.g. the app/package name>
OTEL_RESOURCE_ATTRIBUTES=deployment.environment.name=<env>
OTEL_EXPORTER_OTLP_ENDPOINT=https://ingest.maple.dev
OTEL_EXPORTER_OTLP_HEADERS="Authorization=Bearer <key>"
CREWAI_DISABLE_TELEMETRY=true
CREWAI_TRACING_ENABLED=falseOTLPSpanExporter() with no args appends /v1/traces. If you pass endpoint= in code, it must end in /v1/traces.CREWAI_DISABLE_TELEMETRY=true stops the analytics export to telemetry.crewai.com. CREWAI_TRACING_ENABLED=false stops the AMP uploader and its first-run prompt that waits on stdin at exit. Both must be in the environment before crewai is imported (env file loaded by the process, or os.environ.setdefault(...) at the top of tracing.py if the repo has no env mechanism).Create tracing.py (adapt the module path to the repo layout):
# tracing.py
from openinference.instrumentation import TraceConfig
from openinference.instrumentation.crewai import CrewAIInstrumentor
from openinference.instrumentation.openai import OpenAIInstrumentor
from opentelemetry import trace
from opentelemetry.exporter.otlp.proto.http.trace_exporter import OTLPSpanExporter
from opentelemetry.sdk.trace import SpanProcessor, TracerProvider
from opentelemetry.sdk.trace.export import BatchSpanProcessor
class CrewAIAgentNames(SpanProcessor):
"""Copies each CrewAI agent's role to gen_ai.agent.name, which Maple uses for agent lanes."""
def on_start(self, span, parent_context=None):
# The instrumentor records the role (graph.node.id) just after the agent span starts,
# so name the agent span when its first child starts, while it's still open.
parent = trace.get_current_span(parent_context)
attrs = getattr(parent, "attributes", None) or {}
role = attrs.get("graph.node.id")
if role and "gen_ai.agent.name" not in attrs and parent.is_recording():
parent.set_attribute("gen_ai.agent.name", role)
provider = TracerProvider() # reads OTEL_SERVICE_NAME and OTEL_RESOURCE_ATTRIBUTES
provider.add_span_processor(CrewAIAgentNames())
provider.add_span_processor(BatchSpanProcessor(OTLPSpanExporter()))
trace.set_tracer_provider(provider)
config = TraceConfig(enable_genai_semconv=True)
CrewAIInstrumentor().instrument(tracer_provider=provider, config=config, skip_dep_check=True)
OpenAIInstrumentor().instrument(tracer_provider=provider, config=config, skip_dep_check=True)import tracing at the top of every entry point (web app module, worker, CLI main, main.py of a crewai create project) so instrument() runs before the first kickoff.AnthropicInstrumentor().instrument(...) etc. with the same tracer_provider, config and skip_dep_check=True. Only for SDKs the app uses (Step 0).TracerProvider()/set_tracer_provider, add CrewAIAgentNames() and the exporter to the existing provider, pass it as tracer_provider=.enable_genai_semconv=True on EVERY instrumentor is required. The env var OPENINFERENCE_ENABLE_GENAI_SEMCONV=true is equivalent only if set before instrument(); prefer the code form.skip_dep_check=True: a failed version check makes instrument() skip itself with only an error log.CrewAIAgentNames exactly, and add it BEFORE the BatchSpanProcessor.CrewAI has no conversation id; each kickoff() is its own trace. crew_id (new per Crew object) and crew_key (same for every user of a crew) are NOT conversation ids. Maple reads session.id for CrewAI, and the instrumentors set it only inside using_session:
from crewai import LLM, Agent, Crew, Task
from openinference.instrumentation import using_session
llm = LLM(model="openai/gpt-4o-mini", temperature=0)
def build_crew(text: str, history: str, stream: bool = False) -> Crew:
assistant = Agent(
role="assistant",
goal="Answer the user's questions",
backstory="You are a concise, helpful assistant.",
llm=llm,
tools=[get_weather, calculate],
)
task = Task(
description=f"{text}\n\nConversation so far:\n{history}",
expected_output="A short, direct reply to the user.",
agent=assistant,
name="reply",
)
return Crew(name="support", agents=[assistant], tasks=[task], stream=stream)
def handle_message(conversation_id: str, text: str, history: str) -> str:
with using_session(conversation_id):
return build_crew(text, history).kickoff().rawwith using_session(<conversation id>):. Use the app's stored conversation/chat/thread id. Never a fresh UUID per request, never a constant.with using_session(sid): flow.handle_turn(text, session_id=sid). Use the same id for both. Give the flow class a name = "<snake_name>" attribute: unnamed flows produce Flow_<uuid>.kickoff roots (then Flow_<session id> from the second turn).trace:<id>.Crew a name= (otherwise the root span is Crew_<uuid>.kickoff) and every Task a name=. Put the user's message FIRST in the task description: Maple labels turns with the first line of Current Task: ….akickoff() is NOT instrumented (no crew/agent spans, each model/tool call becomes its own trace). Replace await crew.akickoff(...) with await crew.kickoff_async(...) (same result, runs instrumented kickoff in a thread). Same for Agent.akickoff -> Agent.kickoff in a thread. Tell the user why.Crew(stream=True) calls kickoff twice (an empty stub kickoff span + the real one in a second trace). Wrap each streamed turn:from opentelemetry import trace
tracer = trace.get_tracer("chat")
def stream_message(conversation_id: str, text: str, history: str, send) -> None:
with using_session(conversation_id), tracer.start_as_current_span(
"invoke_agent support",
attributes={
"gen_ai.operation.name": "invoke_agent",
"gen_ai.conversation.id": conversation_id,
},
):
for chunk in build_crew(text, history, stream=True).kickoff():
send(chunk.content)Iterate INSIDE the with. LLM(stream=True) alone (no Crew(stream=True)) needs no wrapper.
@before_tool_call + context.request_human_input(...) need nothing: the hook runs inside the kickoff before the tool span starts (approved call = one tool span; blocked call = no tool span, model gets Tool execution blocked by hook).flow.resume(...) after @human_feedback is not instrumented: wrap it the same way (using_session with the same id + the wrapper span).session.id automatically; give them attributes=dict(get_attributes_from_context()) (from openinference.instrumentation) if you add any beyond the wrapper above.Current Task: … + context) and the reply; agent spans the task and output; tool spans arguments and results. Leave it on unless the user or repo says prompts are sensitive.TraceConfig(enable_genai_semconv=True, hide_inputs=True, hide_outputs=True) passed to every instrumentor (or OPENINFERENCE_HIDE_INPUTS=true / OPENINFERENCE_HIDE_OUTPUTS=true). Narrower: hide_input_text, hide_output_text.crew_tasks (task descriptions, i.e. the user's message), crew_inputs, crew_agents on the <crew>.kickoff span, or flow_inputs on a flow's kickoff span. If prompts must never leave the infrastructure, tell the user to delete those attributes in an OpenTelemetry Collector (attributes processor, action: delete).share_crew=True to get content; it only adds data to CrewAI's analytics.<tool>.run span status ERROR with the message). Do not catch exceptions inside tools to return an error string: the span stays OK and Maple won't count the failure.BaseLLM subclasses or LiteLLM models without function calling use the ReAct text path, which the instrumentor doesn't patch: no tool spans. Tell the user.role; CrewAIAgentNames turns roles into lanes.async_execution=True tasks keep context (siblings under the crew span). Nothing to do.Process.hierarchical: delegated coworker work (Delegate work to coworker / Ask question to coworker tools) runs through un-instrumented Agent.execute_task: its model calls sit inside the tool span, no lane. Known; not fixable here.gen_ai.tool.call.id on tool spans.TracerProvider flushes on normal interpreter exit (atexit). That covers servers and CLIs that exit normally, including crewai run.provider.force_flush() in a finally before returning.os._exit, and notebooks: provider.force_flush() after each run; provider.shutdown() at the very end.Run one real conversation (2-3 messages, same conversation id, at least one tool call), and one message in a second conversation. No scriptable entry point (server, UI, REPL only): write a small driver for this run (one conversation id, 2+ turns, one tool call, flush before exit).
Without Maple access (MAPLE_TEST, no MCP): the run must exit with no export errors on stderr (Failed to export, 401 lines) AND a local exporter (provider.add_span_processor(SimpleSpanProcessor(ConsoleSpanExporter())) in a scratch run) must show <crew>.kickoff, model and <tool>.run spans carrying session.id. Silence alone proves nothing (no spans is silent too). Say so and list what the user should check in Maple. With the Maple MCP: list_agent_sessions with search=<conversation id> returns one row.
Otherwise check in Maple Agent Sessions (https://app.maple.dev/agent-sessions, EU app.eu.maple.dev), filtered to the service name:
trace:<id> sessions.openinference-openai; that's fine).<crew name>.kickoff (or <flow name>.kickoff, or your invoke_agent wrapper when streaming). No empty extra turns.Current Task: …, replies).ChatCompletion for the OpenAI instrumentor) have a model and non-zero input/output tokens, including streamed calls. Each call appears once.<tool>.run with results; a tool that raised is counted as failed and nothing else is.<role>.<task>._execute_core spans carry gen_ai.agent.name).llm.cost.total). Expected.crewai.telemetry, no coding_agent attribute (telemetry is off).Edge cases:
gen_ai.usage.input_tokens/output_tokens (plus cached and reasoning tokens when reported); crew and agent spans carry none, so nothing double-counts. CrewAI's OpenAI provider always requests stream_options={"include_usage": True} when streaming, so streamed calls keep tokens.anthropic/claude-haiku-4.5 behind OpenRouter); provider = the SDK used, so every OpenRouter model shows openai.memory=True / planning=True add real, billed model calls (memory analysis, embeddings, planning agent); they appear in the session. Expected.<flow name>.kickoff root, one <flow name>.<method> span per @start/@listen/@router method, crews and Agent.kickoff() nested inside. Conversational flow turns show <flow>.route_conversation and <flow>.converse_turn under the kickoff.If sessions are split per message: using_session missing or id changing. No model spans/tokens: wrong or missing SDK instrumentor. Every call its own trace: akickoff. Nothing arrives: exporter endpoint/header wrong, .env loaded after tracing.py, OTEL_SDK_DISABLED=true, or process exited without flushing. 401 ingest_unauthorized / "Invalid ingest key" with a key you trust: keys are region-bound, so it likely belongs to the other region; try the other endpoint.
litellm.callbacks=["otel"]) on the same calls: duplicate model spans and tokens.49cb661
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.