Use this skill when the user asks to design, audit, or harden Claude Platform Messages API context budgets, token counting, prompt caching, cache diagnostics, context editing, server-side compaction, context-window overflow policy, or mid-conversation system-message modes around a fixed effort setting. Do not use it for tool dispatch correlation, durable memory systems, generic context-writing advice, executor implementation, or Bash ownership.
Own the client policy that measures, preserves, diagnoses, reduces, and bounds the context sent to the Claude Platform Messages API. [DOC]
Load references/official-source-map.md before asserting API fields. The versioned audit and normalized local corpus are the evidence base; details not covered there remain coverage_gap. [CONFIG]
Use this skill for token preflight, cache placement and accounting, cache-miss diagnosis, selective server-side context edits, compaction continuity, context-window overflow, web/tool-search context pressure, and bounded mid-conversation system-message modes. [DOC]
Route elsewhere when the primary work is:
claude-api-tool-runtime: tool_use / tool_result correlation, dispatch, stop routing, strict input validation, or partial tool JSON. [CONFIG]tool-use-design: tool names, descriptions, schemas, and catalog ergonomics. [CONFIG]headless-sdk-automation: Claude Code CLI, claude -p, CI, or Agent SDK sessions. [CONFIG]claude-api-server-tools: tool search, web search/fetch, server execution, and server continuation; this skill owns only their context/cache implications. [CONFIG]claude-api-files-documents: Files API and PDF/document input contracts; this skill owns only context fit and reuse. [CONFIG]claude-api-client-tools: application-owned memory/bash/editor/computer handlers; durable state and sandbox policy stay there. [CONFIG]This skill does not own durable memory architectures, generic prompting curricula, or the multi-agent executor shown in the effort example. It owns only the Messages API context controls around those workloads. [CONFIG]
| Surface | Changes context occupancy? | Primary evidence |
|---|---|---|
| Prompt caching | No; it reuses an identical prefix for latency/cost accounting. | prompt-caching, tool-use-with-prompt-caching [DOC] |
| Cache diagnostics | No; it compares request fingerprints and reports the earliest divergence. | cache-diagnostics [DOC] |
| Context editing | Yes; it removes selected old tool-result or thinking content. | context-editing [DOC] |
| Compaction | Yes; it replaces older history with a returned summary block. | compaction [DOC] |
| Window limits | Bounds input plus generated output; cached tokens still count. | context-windows, token-counting [DOC] |
coverage_gap. [DOC][INFERENCIA]references/source-map.json; do not infer model limits, beta headers, enums, or TTL behavior from names alone. [CONFIG]scripts/validate_context_policy.py, then run scripts/check.sh. [CÓDIGO]max_tokens, provider surface, active beta headers, and SDK/API version. [SUPUESTO]tools, system, messages, images/documents, thinking mode, context-management edits, and cache controls. [DOC]usage fields when diagnosing caching. Counts are estimates and must be recomputed for the target model. [DOC]references/context-strategy.md; distinguish web-search result pressure from tool-search definition pressure, and do not treat either as a generic cache problem. [DOC]references/window-and-token-accounting.md. Cached reads and writes remain part of total input occupancy. [DOC]references/prompt-caching-policy.md: stable prefix, tools -> system -> messages invalidation hierarchy, breakpoint/TTL ordering, and append-only history. [DOC]references/cache-diagnostics-policy.md. Combine diagnostics with usage.cache_read_input_tokens; neither alone proves a cache hit. [DOC]references/context-reduction-policy.md: selective deletion for stale blocks, compaction for summarized continuity. Preserve returned compaction blocks and aggregate usage.iterations. [DOC]references/mid-conversation-controls.md for trusted operator messages and finite loop bounds. For the effort example, keep EFFORT = "xhigh" fixed in every request and toggle only orchestration mode through appended system messages. Never place a system message between tool_use and its tool_result. [DOC]templates/context-management-report.md, validate its policy JSON offline, and have agents/verifier.md review evidence and failure paths independently of agents/producer.md. [CONFIG]input_tokens + cache_read_input_tokens + cache_creation_input_tokens; never subtract cached tokens from the window budget. [DOC]data-privacy-governance. [DOC][CONFIG]diagnostics: null, pending comparison, previous_message_not_found, and unavailable are not cache-hit evidence. [DOC]allowed_callers path. Direct-only models require allowed_callers: ["direct"], and continuation must preserve returned result blocks and encrypted_content. [DOC]tool_reference continuity, and a cache breakpoint on a non-deferred tool. [DOC]compact_20260112, and the minimum trigger; return the compaction block on later requests, handle stop_reason: compaction when pausing, consume its single full streaming delta, and count all sampling through usage.iterations. [DOC]EFFORT = "xhigh" constant for the main loop and subagents while system messages toggle orchestration mode. [DOC]tool-permission-policy. [DOC][CONFIG]claude-api-files-documents. [DOC][CONFIG]coverage_gap and reject production approval. [CONFIG]coverage_gap. [CÓDIGO]references/official-source-map.md and references/source-map.json - canonical source IDs, URLs, hashes, and local evidence routes.references/context-policy-contract.md - normalized policy shape and fail-closed issue codes.references/context-strategy.md - pressure-to-mechanism routing.references/prompt-caching-policy.md - breakpoints, invalidation, TTL, and tool-prefix rules.references/cache-diagnostics-policy.md - diagnostic states and usage matrix.references/context-reduction-policy.md - context editing versus compaction.references/window-and-token-accounting.md - preflight and overflow math.references/mid-conversation-controls.md - system-message placement and effort bounds.assets/context-management-checklist.md - producer/verifier gate.scripts/validate_context_policy.py - importable, stdlib-only validator.scripts/check.sh - deterministic packet, source, fixture, and unittest gate.Capas del packet, cargables bajo demanda (disciplina ICM: una capa por vez, nunca todas juntas): references/ guías de profundidad (cargar UNA por etapa) · knowledge/ cuerpo de conocimiento · prompts/ prompts listos · examples/ salida de ejemplo · agents/ subagentes del packet · templates/ plantilla de output · scripts/ automatización local · assets/ recursos estáticos.
e8f986b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.