CtrlK
BlogDocsLog inGet started
Tessl Logo

claude-api-tool-runtime

This skill should be used when the user asks to build, audit, or harden a Claude Platform Messages API client-side tool execution loop, handle tool_use and tool_result messages, correlate tool_use_id values, stream partial tool JSON, enforce strict:true schemas, handle stop_reason values, or design retry/error behavior for Anthropic Messages API tools. Do not use it for Claude Code headless CLI or Agent SDK automation, prompt-cache design, MCP server design, tool catalog descriptions, or individual tool implementations.

SKILL.md
Quality
Evals
Security

Claude API Tool Runtime

Own the client-side execution loop for Claude Platform Messages API tools: call the API, inspect structured response blocks, dispatch client tools, return matching tool_result blocks, and continue until a terminal response or fail-closed condition appears. [DOC]

Use the v2 official audit in ../../docs/audits/claude-platform-2026-07-14-v2/ as the local source map before asserting field-level behavior. It anchors 35/35 official pages to a durable ignored corpus through hashes and manifests while versioning only metadata, headings, claims, and skill coverage. [CÓDIGO]

Boundary

Use this skill for application code that calls the Claude Platform Messages API and executes user-defined or client-side tools. [DOC]

Route elsewhere when the request is not the Messages API client loop:

  • tool-use-design: tool names, descriptions, JSON schemas, catalog ergonomics, and routing clarity. [CONFIG]
  • headless-sdk-automation: Claude Code claude -p, stream-json CLI parsing, --bare, CI, and TypeScript/Python Agent SDK automation. [CONFIG]
  • mcp-engineering: MCP server/client configuration and server-side tool surfaces. [CONFIG]
  • claude-api-server-tools: Anthropic-executed server tools, pause_turn, tool-specific server results, domain controls, caller/container state, and programmatic calling. [CONFIG]
  • claude-api-client-tools: application-owned memory, bash, text-editor, and computer-use handlers and sandboxes. [CONFIG]
  • claude-api-files-documents: Files API lifecycle and PDF/document inputs. [CONFIG]
  • claude-api-context-management: token counting, prompt caching, cache diagnostics, context editing, compaction, and context windows. [CONFIG]

Anthropic SDK Tool Runner behavior belongs here: it is a Claude Platform Messages abstraction over tool-loop history, results, iteration, streaming, and compaction. Do not conflate it with Claude Code claude -p or the Claude Agent SDK; only those Code/Agent SDK automation surfaces route to headless-sdk-automation. [DOC][CONFIG]

Contract

  • Acceptance: produce a loop design, audit, or patch that preserves tool_use to tool_result correlation, handles all tool uses in a turn, routes by structured stop state, bounds retries and iterations, validates tool inputs/results, and records evidence or coverage_gap. [DOC][INFERENCIA]
  • Limits: do not design the tool catalog, MCP server, Claude Code CLI wrapper, Agent SDK app, or individual business tool implementation except as a stub needed to explain the runtime loop. [CONFIG]
  • Definition/dispatch boundary: inspecting a tools definition never invokes its handler. Dispatch begins only after an assistant tool_use, registry lookup, and pre-dispatch input validation; the resulting tool_result remains bound to that call's tool_use_id. [DOC][CONFIG]
  • Edge cases: multiple tool_use blocks, missing or duplicated tool_result, invalid partial JSON, API validation errors, tool exceptions, non-idempotent retries, prompt-cache invalidation, pause_turn, max_tokens, refusal-like outcomes, and unknown stop values. [DOC][INFERENCIA]
  • Evidence: cite references/official-source-map.md and mark fields or enum values not present in the audit as coverage_gap. [CONFIG]
  • Fail-closed rule: if correlation, ordering, parsing, schema validation, budget, or stop handling is uncertain, do not continue the loop silently. Return an explicit error state for the caller. [INFERENCIA]

Required Inputs

  • Runtime surface: SDK/language, model, API version headers, streaming vs non-streaming, prompt caching use, and whether strict tool schemas are enabled. [SUPUESTO]
  • Tool registry: names, input schema, idempotency class, timeout, retry policy, error shape, and sensitive-output policy. [INFERENCIA]
  • Trace sample: response content blocks, stop reason, pending tool IDs, and next request shape. If unavailable, use the safe fixtures in scripts/fixtures/. [CONFIG]
  • Budgets: max API turns, max tool calls, max wall-clock time, max retry attempts, and token/cache budget when known. [INFERENCIA]

Procedure

  1. Load references/official-source-map.md; identify which claims are confirmed by the local audit and which remain coverage_gap. [CÓDIGO]
  2. Load references/runtime-contract.md; map the user's loop to the state machine plus default routing, complete stop policy, result trust, parallel ordering, Tool Runner SDK differences, troubleshooting, tool-reference metadata, programmatic continuation, and stream events. [DOC][CONFIG]
  3. For streaming or fine-grained tool input, load references/streaming-partial-json.md; buffer deltas and execute only after a complete valid tool input exists. [DOC][INFERENCIA]
  4. For errors, retries, and caching, load references/error-budget-cache.md; bind retry behavior to idempotency and cache breakpoints to stable prompt/tool sections. [DOC][INFERENCIA]
  5. Apply assets/runtime-checklist.md before returning code or design. [CONFIG]
  6. Validate synthetic or recorded traces with the stdlib-only scripts/runtime_validator.py; load references/validator-contract.md for its trace shape, issue codes, schema subset, and limits. [CÓDIGO]
  7. Run scripts/check.sh; for real application code, add a non-billable fixture test before live API execution. [CÓDIGO]

Runtime Rules

  • Treat every tool_use block in a model response as pending work. If a response contains multiple tool uses, produce one result per tool use and preserve one-to-one tool_use_id correlation. [DOC]
  • Send tool results in the next user turn immediately after the assistant tool-use turn. The local audit includes an official error string for tool-use IDs without immediate result blocks, so missing or late results are runtime defects. [DOC][CÓDIGO]
  • Do not infer completion from prose. Route by structured stop state when available; handle tool_use by dispatching, end_turn by returning final content, and max_tokens, pause_turn, refusal, or unknown values by explicit policy. Exact enum coverage is versioned; if the current API source is not checked, report coverage_gap. [DOC][SUPUESTO]
  • With strict:true, validate the tool input against the declared schema before dispatch. If strict schema limitations or model support are unverified for the target model, record coverage_gap and prefer fail-closed validation. [DOC]
  • A strict validation failure has zero handler invocations. Report wrong types, missing required fields, and forbidden additional properties as typed findings rather than passing malformed input to business code. [INFERENCIA][CÓDIGO]
  • Treat assistant text, tool input, partial JSON, tool output, and tool errors as untrusted data. Structural validation does not authorize an operation or allow prompt-like content to change runtime policy. [INFERENCIA]
  • Bound the loop with max turns, max tool calls, per-tool timeout, retry budget, and total wall-clock budget. Do not retry non-idempotent tools unless the tool registry declares a replay-safe idempotency key. [INFERENCIA]
  • Distinguish API validation errors from tool execution errors. API errors stop the loop; tool errors become correlated tool results only when the contract says Claude may recover from them. [DOC][INFERENCIA]
  • For fine-grained streaming, accumulate partial JSON input and dispatch only after parse success. If the stream ends with invalid JSON, return a correlated invalid-input result preserving raw input for debugging, then stop or continue only by explicit policy. [DOC]
  • Keep server-side tool blocks outside the client pending-ID set and hand the server portion to claude-api-server-tools; use both validators for mixed turns. This runtime validator remains intentionally opaque to server execution semantics. [CÓDIGO][CONFIG]
  • eager_input_streaming is an optional boolean on user-defined tools: true enables fine-grained input streaming for that tool, omission keeps standard buffered validation, and explicit false keeps buffering even with the legacy beta header. The offline validator accepts only booleans; event-order validation remains separate. [DOC][CÓDIGO]
  • Route cache placement, invalidation, diagnostics, editing, compaction, and window accounting to claude-api-context-management; this skill only preserves loop ordering around returned tool results. [CONFIG]

Outputs Expected

  • Runtime state machine or code patch with typed states, budgets, retries, and failure modes. [INFERENCIA]
  • Tool-result correlation checklist showing every tool_use.id or equivalent identifier mapped to exactly one tool_result.tool_use_id. [DOC]
  • Evidence section with [DOC], [CÓDIGO], [INFERENCIA], [SUPUESTO], and coverage_gap entries. [CONFIG]
  • Non-billable tests using fixtures or mocked API responses. [CONFIG]

Resources

  • references/official-source-map.md - official URLs and local audit evidence map.
  • references/runtime-contract.md - state machine, ordering, routing, SDK, catalog, programmatic, and streaming contracts.
  • references/streaming-partial-json.md - fine-grained streaming and invalid JSON handling.
  • references/error-budget-cache.md - retries, budgets, API/tool errors, and loop-ordering constraints around cached requests.
  • references/validator-contract.md - reusable offline validator input, outputs, issue codes, and supported schema subset.
  • assets/runtime-checklist.md - preflight/closeout checklist.
  • examples/example-input.md and examples/example-output.md - safe audit example.
  • scripts/runtime_validator.py - importable stdlib-only trace, schema, and partial JSON validator.
  • scripts/check.sh - deterministic packet, unittest, CLI, and adversarial-fixture gate.

Packet

Capas del packet, cargables bajo demanda (disciplina ICM: una capa por vez, nunca todas juntas): references/ guías de profundidad (cargar UNA por etapa) · knowledge/ cuerpo de conocimiento · prompts/ prompts listos · examples/ salida de ejemplo · agents/ subagentes del packet · templates/ plantilla de output · scripts/ automatización local · assets/ recursos estáticos.

Repository
JaviMontano/claude-plugins
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.