CtrlK
BlogDocsLog inGet started
Tessl Logo

claude-api-context-management

Use this skill when the user asks to design, audit, or harden Claude Platform Messages API context budgets, token counting, prompt caching, cache diagnostics, context editing, server-side compaction, context-window overflow policy, or mid-conversation system-message modes around a fixed effort setting. Do not use it for tool dispatch correlation, durable memory systems, generic context-writing advice, executor implementation, or Bash ownership.

SKILL.md
Quality
Evals
Security

Claude API Context Management

Own the client policy that measures, preserves, diagnoses, reduces, and bounds the context sent to the Claude Platform Messages API. [DOC]

Load references/official-source-map.md before asserting API fields. The versioned audit and normalized local corpus are the evidence base; details not covered there remain coverage_gap. [CONFIG]

Boundary

Use this skill for token preflight, cache placement and accounting, cache-miss diagnosis, selective server-side context edits, compaction continuity, context-window overflow, web/tool-search context pressure, and bounded mid-conversation system-message modes. [DOC]

Route elsewhere when the primary work is:

  • claude-api-tool-runtime: tool_use / tool_result correlation, dispatch, stop routing, strict input validation, or partial tool JSON. [CONFIG]
  • tool-use-design: tool names, descriptions, schemas, and catalog ergonomics. [CONFIG]
  • headless-sdk-automation: Claude Code CLI, claude -p, CI, or Agent SDK sessions. [CONFIG]
  • claude-api-server-tools: tool search, web search/fetch, server execution, and server continuation; this skill owns only their context/cache implications. [CONFIG]
  • claude-api-files-documents: Files API and PDF/document input contracts; this skill owns only context fit and reuse. [CONFIG]
  • claude-api-client-tools: application-owned memory/bash/editor/computer handlers; durable state and sandbox policy stay there. [CONFIG]

This skill does not own durable memory architectures, generic prompting curricula, or the multi-agent executor shown in the effort example. It owns only the Messages API context controls around those workloads. [CONFIG]

Distinctions

SurfaceChanges context occupancy?Primary evidence
Prompt cachingNo; it reuses an identical prefix for latency/cost accounting.prompt-caching, tool-use-with-prompt-caching [DOC]
Cache diagnosticsNo; it compares request fingerprints and reports the earliest divergence.cache-diagnostics [DOC]
Context editingYes; it removes selected old tool-result or thinking content.context-editing [DOC]
CompactionYes; it replaces older history with a returned summary block.compaction [DOC]
Window limitsBounds input plus generated output; cached tokens still count.context-windows, token-counting [DOC]

Contract

  • Acceptance: return a source-grounded context policy with a preflight count, window reserve, explicit cache strategy, diagnostic interpretation, reduction strategy, continuity rules, fixed request effort when the example requires it, finite loop bounds, and residual coverage_gap. [DOC][INFERENCIA]
  • Fail closed: reject unknown policy fields, unverified window capacity, projected overflow, ambiguous diagnostic states, unsafe system-message placement, unsupported edit combinations, lost compaction blocks, or unbounded effort loops. [CONFIG]
  • Evidence: use the ten source IDs in references/source-map.json; do not infer model limits, beta headers, enums, or TTL behavior from names alone. [CONFIG]
  • Verification: validate policy JSON with scripts/validate_context_policy.py, then run scripts/check.sh. [CÓDIGO]

Required Inputs

  • Model ID, verified context-window limit, requested max_tokens, provider surface, active beta headers, and SDK/API version. [SUPUESTO]
  • Structured request ingredients: tools, system, messages, images/documents, thinking mode, context-management edits, and cache controls. [DOC]
  • Token-count estimate plus the latest response usage fields when diagnosing caching. Counts are estimates and must be recomputed for the target model. [DOC]
  • Retention requirements: which tool results, thinking turns, recent messages, and operator instructions must remain verbatim. [SUPUESTO]
  • Finite limits for output reserve, safety margin, turns, concurrency, subtasks, timeouts, and compaction events. [INFERENCIA]

Procedure

  1. Classify pressure with references/context-strategy.md; distinguish web-search result pressure from tool-search definition pressure, and do not treat either as a generic cache problem. [DOC]
  2. Count the same structured request and target model; apply references/window-and-token-accounting.md. Cached reads and writes remain part of total input occupancy. [DOC]
  3. If caching is useful, apply references/prompt-caching-policy.md: stable prefix, tools -> system -> messages invalidation hierarchy, breakpoint/TTL ordering, and append-only history. [DOC]
  4. If a hit is unexpectedly absent, apply references/cache-diagnostics-policy.md. Combine diagnostics with usage.cache_read_input_tokens; neither alone proves a cache hit. [DOC]
  5. Choose exactly the needed reduction mechanism with references/context-reduction-policy.md: selective deletion for stale blocks, compaction for summarized continuity. Preserve returned compaction blocks and aggregate usage.iterations. [DOC]
  6. Apply references/mid-conversation-controls.md for trusted operator messages and finite loop bounds. For the effort example, keep EFFORT = "xhigh" fixed in every request and toggle only orchestration mode through appended system messages. Never place a system message between tool_use and its tool_result. [DOC]
  7. Materialize templates/context-management-report.md, validate its policy JSON offline, and have agents/verifier.md review evidence and failure paths independently of agents/producer.md. [CONFIG]

Hard Rules

  • Prompt caching changes reuse and billing, not the number of tokens occupying the context window. [DOC]
  • Compute observed total input as input_tokens + cache_read_input_tokens + cache_creation_input_tokens; never subtract cached tokens from the window budget. [DOC]
  • A token-count response is an estimate, not a billing guarantee or exact future message count. [DOC]
  • Extended-thinking accounting is model-specific; do not generalize prior-thinking retention or exclusion across model families when official pages differ. [DOC][CONFIG]
  • Cache diagnostics expose bounded hash/token fingerprints rather than raw prompts. Preserve documented ZDR, retention, and workspace/organization isolation qualifiers, and route broader policy conclusions to data-privacy-governance. [DOC][CONFIG]
  • diagnostics: null, pending comparison, previous_message_not_found, and unavailable are not cache-hit evidence. [DOC]
  • Context editing invalidates affected cached prefixes; require a deliberate removal threshold when caching and tool-result clearing coexist. [DOC][INFERENCIA]
  • Writing essential results to memory before clearing them is an explicit client-tool handoff, not storage provided by context editing itself. [DOC][CONFIG]
  • Web search basic results enter the model context; dynamic filtering on supported versions reduces that pressure through the documented allowed_callers path. Direct-only models require allowed_callers: ["direct"], and continuation must preserve returned result blocks and encrypted_content. [DOC]
  • Tool search keeps deferred definitions out of the initial model context, not out of the request catalog. Preserve the complete catalog, server result blocks, API-expanded tool_reference continuity, and a cache breakpoint on a non-deferred tool. [DOC]
  • Compaction is not deletion: validate the beta, supported request model, compact_20260112, and the minimum trigger; return the compaction block on later requests, handle stop_reason: compaction when pausing, consume its single full streaming delta, and count all sampling through usage.iterations. [DOC]
  • Mid-conversation system content must be trusted operator context, append-only, correctly placed, and never raw tool/web/retrieval output. [DOC]
  • The mid-conversation effort page does not dynamically change effort: it holds EFFORT = "xhigh" constant for the main loop and subagents while system messages toggle orchestration mode. [DOC]
  • The example's application-owned Bash handler executes model-written commands locally without a sandbox. Treat that only as a risk boundary; this skill does not own the executor or Bash and does not route the Platform Bash capability to tool-permission-policy. [DOC][CONFIG]
  • Media-page limits are qualified by the verified context window and remain separate from request-size/PDF payload limits owned by claude-api-files-documents. [DOC][CONFIG]
  • If the target model/provider/version is not verified against the active source set, emit coverage_gap and reject production approval. [CONFIG]

Outputs Expected

  • Boundary decision and selected context strategy. [CONFIG]
  • Token/window ledger with projected occupancy and remaining margin. [DOC][INFERENCIA]
  • Separate caching, diagnostics, editing, compaction, and overflow policies. [CONFIG]
  • Mid-conversation placement, fixed-effort, mode-toggle, and finite-loop policy. [CONFIG]
  • Offline validator report, verifier decision, and residual coverage_gap. [CÓDIGO]

Resources

  • references/official-source-map.md and references/source-map.json - canonical source IDs, URLs, hashes, and local evidence routes.
  • references/context-policy-contract.md - normalized policy shape and fail-closed issue codes.
  • references/context-strategy.md - pressure-to-mechanism routing.
  • references/prompt-caching-policy.md - breakpoints, invalidation, TTL, and tool-prefix rules.
  • references/cache-diagnostics-policy.md - diagnostic states and usage matrix.
  • references/context-reduction-policy.md - context editing versus compaction.
  • references/window-and-token-accounting.md - preflight and overflow math.
  • references/mid-conversation-controls.md - system-message placement and effort bounds.
  • assets/context-management-checklist.md - producer/verifier gate.
  • scripts/validate_context_policy.py - importable, stdlib-only validator.
  • scripts/check.sh - deterministic packet, source, fixture, and unittest gate.

Packet

Capas del packet, cargables bajo demanda (disciplina ICM: una capa por vez, nunca todas juntas): references/ guías de profundidad (cargar UNA por etapa) · knowledge/ cuerpo de conocimiento · prompts/ prompts listos · examples/ salida de ejemplo · agents/ subagentes del packet · templates/ plantilla de output · scripts/ automatización local · assets/ recursos estáticos.

Repository
JaviMontano/claude-plugins
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.