Advisory-only Kafka tuning for Confluent Cloud or Confluent Platform by priority: latency, throughput, availability, or durability. Use when asked to size a planned topic or tune an existing cluster, topic, producer, or consumer for a quantified target such as lower p99 latency or no lost acknowledged writes. Inspects configuration read-only and recommends changes, trade-offs, and validation; the administrator applies them. Do NOT trigger for: WarpStream or Apache Kafka OSS-only clusters; Flink SQL or UDF authoring, tuning, or debugging (use confluent-cloud-flink-sql or flink-udf); Kafka Connect tuning; new Java/Python client implementation (use developing-kafka-java-client or developing-kafka-python-client); Kafka Streams/KStream/KTable development (use kafka-streams-programming); or CDC, Tableflow, or Iceberg pipelines (use confluent-cloud-cdc-tableflow).
70
88%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Turns a workload's priority (latency, throughput, availability, or durability) into a concrete, platform-appropriate recommendation: which Kafka settings to change, from what to what, why, what it trades against, and how the administrator should validate it afterward. Targets Confluent Cloud and Confluent Platform only. Kafka only — Flink (including Flink SQL job tuning) is out of scope; see the "Do NOT trigger" clause in the description.
These four priorities are not independent — every recommendation trades against at least one other axis. The core job of this skill is to make that trade-off explicit, not to pretend a free lunch exists.
Advisory-only — this skill never changes anything. It may inspect configuration and runtime state (read-only) and propose changes. It must never apply changes, run write/produce operations, or otherwise mutate a broker, topic, or client — the administrator or developer is solely responsible for reviewing and applying any recommendation. Keep recommendation and execution strictly separate: present the proposed change, do not ask for permission to apply it, do not apply it, and never state or imply that a change has been made. This skill ships no executable scripts; the only actions it takes are read-only config/metric inspection via the Confluent CLI, an MCP server, or the Kafka REST v3 / REST Proxy v3 and Cloud Metrics APIs.
| Situation | Read this |
|---|---|
| User is targeting Confluent Cloud | references/confluent-cloud.md |
| User is targeting Confluent Platform (self-managed) | references/confluent-platform.md |
| Any config knob lookup by priority (producer/consumer/topic/broker) | references/tuning-parameters.md |
| Verifying a Platform change, or choosing which broker JMX metric proves it worked (Step 7) | references/monitoring.md |
| Verifying a Cloud change, or choosing which metric proves it worked (Step 7) | references/confluent-cloud.md for server-side metrics; use references/monitoring.md only for the client-metrics sections |
Do not read all of these upfront — pull in only what the current step needs.
references/tuning-parameters.mdAsk, in one pass:
If the user only says which axis matters and nothing else, proceed with the interview above before touching any reference file — the recommendation depends on the quantified target, not just the axis name.
Ask which platform this workload runs on: Confluent Cloud or Confluent Platform (self-managed). This skill does not cover WarpStream or unmanaged Apache Kafka — say so and stop if the user names one of those. It also does not cover Kafka Connect (connector or worker tuning) — if the request is about a connector or Connect worker, say it's out of scope and stop.
Flink is a hard stop. If the request mentions Flink at all — a Flink SQL job, a Flink pipeline,
parallelism, checkpointing, state backend, a lagging/slow Flink statement, a UDF — this skill does
not apply. Say it's out of scope and route the user to confluent-cloud-flink-sql (Flink SQL) or
flink-udf (Flink UDFs), then stop. Do not proceed into the workflow, and do not offer
to tune the underlying Kafka topic/producer/consumer as a fallback or consolation — declining is the
whole response. (If the user later comes back with a request that is solely about Kafka topic or
client configuration, with no Flink tuning ask, treat that as a fresh in-scope request.)
Every out-of-scope handoff follows that same rule. Whenever this skill declines — new client
application code (developing-kafka-java-client, developing-kafka-python-client), Kafka Streams
(kafka-streams-programming), Kafka Connect, CDC/Tableflow (confluent-cloud-cdc-tableflow),
WarpStream, or Apache Kafka OSS — name the right skill and stop. Do not attach tuning
recommendations, default values, or "worth deciding up front" configuration advice to the handoff.
Offering acks, partition-count, or any other setting alongside a decline is still a tuning
recommendation, and it is exactly what makes the handoff fail. Declining is the whole response.
Read the matching reference file now:
Before recommending anything, capture what's actually configured today — recommendations phrased
as diffs are easier to reason about and to confirm than absolute values. This is a read-only
inspection (a describeConfigs); it changes nothing. Use whichever of these the environment
already has:
For a planned topic that does not exist, do not attempt to describe it. Record its topic config as "not created", then inspect only the cluster defaults and capacity limits needed to size it. In Step 5, present proposed initial values rather than a current → recommended diff for that topic.
For broker/cluster-level baselines (Platform only — Cloud does not expose these), use the CLI/REST
equivalent of describeConfigs against the broker resource. Capture the values you'll be
proposing to change, so Step 5 can be written as a clean current → recommended diff.
If a required client or broker baseline cannot be inspected or the user has not supplied it, do not infer a remembered default and present it as current state. Ask for the missing effective values, or limit the response to diagnostic directions and clearly labeled candidate changes until the baseline is available.
Two rules govern every value the recommendation reports:
is_read_only=true),
never propose changing, removing, or overriding it. Note that the platform manages it and move on.Open references/tuning-parameters.md and read the row for
each config surface relevant to the workload (producer, consumer, topic/replication,
broker/cluster). It gives, per priority, the recommended value and the one-line reason. Do not
copy every row blindly — reconcile with the quantified target from Step 1. For example, an
"availability" priority with only 2 brokers available changes what min.insync.replicas can
safely be set to; say so explicitly rather than recommending a value the cluster can't sustain.
Produce a recommendation, not a request to act. Do not ask "should I apply this?" and do not apply anything — the administrator owns execution. Use this format:
Tuning recommendation — priority: <latency|throughput|availability|durability>
Target: <the quantified goal from Step 1>
Platform: <Confluent Cloud | Confluent Platform>
Proposed changes (current → recommended):
<config.key> <current value> → <recommended value> (why: <one line>)
<config.key> <current value> → <recommended value> (why: <one line>)
...
For `replication.factor` changes, follow the topic-creation and reassignment guidance in
[references/tuning-parameters.md](references/tuning-parameters.md).
Never recommend deleting and recreating an existing topic solely to change its replication factor;
on Confluent Platform, preserve the topic and use an explicit replica reassignment after confirming
broker and rack capacity.
Expected impact: <what should measurably improve, tied to the Step 1 target>
Trade-off / risk: <what gets worse on another axis, stated plainly — e.g.
"min.insync.replicas 3→2 improves availability during a single-broker outage but
means a write can be acknowledged with one fewer copy durably stored">
Prerequisites: <anything that must be true/checked first — e.g. broker.rack configured,
cluster tier supports multi-zone, enough brokers to sustain the new min.insync.replicas>
How to apply (for the administrator to run — this skill does not run these):
1. <description>
Command/tool: <operation and current documentation link; include exact syntax only after
verifying it against the installed CLI, discovered MCP tool, or current API documentation>
2. ...
Validation (Step 7) — how the administrator confirms it worked after applying:
Server-side metric: <a metric name confirmed available on this platform — a Platform JMX MBean,
or a Cloud metric returned by live descriptor discovery. When the target is Cloud and
descriptors are unavailable, name no server-side metric here and emit exactly this line:
Server-side metric validation: pending descriptor access>
Client/application signal: <an accessible client metric or application-level SLI>
Benchmark (optional, administrator-run): <the before/after measurement from Step 7>Present the "How to apply" commands as reference material the administrator executes on their own authority — never as something this skill will run, and never phrased as a request for permission to run them. If the user pushes back or asks for a different balance, revise the recommendation and re-present it.
This skill does not apply the change. It stops at the recommendation. The administrator or developer reviews it and applies it themselves, on their own authority, using whatever tooling they already have. The "How to apply" block in Step 5 lists the exact operations for them to run — so they can review and execute without guesswork — using one of:
When the change is disruptive (e.g. a broker restart, or a broker-level default that affects every topic on the cluster), say so and recommend the administrator apply it one broker at a time and confirm cluster health between brokers — as guidance in the recommendation, not as an action this skill takes. Never run any of these operations yourself, and never say a change "has been applied"; at most, note what the administrator will observe once they apply it.
Hand the administrator a way to confirm the change actually moved the target metric — don't let them just trust that the config took effect. These are steps they run after applying; this skill does not run them (a benchmark produces and consumes records, which is a write/mutation, and is out of scope for an advisory-only skill).
Two parts to the validation you recommend:
UnderReplicatedPartitions for availability/durability, the request-latency breakdown for
latency, records-lag-max for consumer throughput. On Confluent Platform these are JMX MBeans;
on Confluent Cloud, use the available Metrics API equivalents and client metrics rather than
assuming every broker JMX MBean is exposed. When giving Platform JMX names, label them as
Platform-only and include this Cloud distinction explicitly, even if the current target is
Platform. This is read-only and is the primary validation.Cloud validation hard stops: validation for Confluent Cloud must never include stopping, restarting, killing, isolating, or otherwise fault-injecting a managed broker — not as an administrator step, optional test, hypothetical procedure, or future recommendation. Do not describe such a procedure in the response. Validate only through the documented cluster-type guarantee, effective topic and producer configuration, accessible client/application signals, and server-side metrics confirmed by live Metrics API descriptor discovery. Never reuse a Platform JMX MBean name as a Cloud metric name unless discovery returns that exact name. If cluster-type metadata, Metrics API credentials, or descriptors are unavailable, explicitly mark those parts of validation pending; do not replace them with broker manipulation or inferred metric names.
An optional before/after benchmark the administrator runs themselves. If they want a measured before/after, recommend they run it with their own tooling — this skill ships nothing and runs nothing (benchmarking produces/consumes records, a write, which is the administrator's to do). Point them at whatever they already have:
kafka-producer-perf-test / kafka-consumer-perf-test (ship with Kafka/Confluent Platform).Recommend running it once before applying and once after, comparing p99 (tail) latency (not
just the average — SLAs are written against the tail) and achieved throughput, measuring
end-to-end latency as producer send() → consumer poll().
Frame the result against the Step 1 target and restate the trade-off from Step 5, so the administrator knows what to expect to improve and what to expect to get worse. There is no one-size-fits-all config, so the administrator's own before/after measurement, not the recommendation table, is the proof.
If, after the administrator applies and validates a recommendation, the target still isn't met, re-inspect the new baseline (Step 3) and refine the recommendation (Step 4). Common reasons a single pass isn't enough: partition count is a bottleneck (throughput), client-side batching settings weren't also changed to match the new broker/topic config, or the workload is bound by something outside Kafka entirely (serialization cost, network path, downstream consumer processing time) — say so rather than continuing to recommend Kafka knobs that won't help.
.env file contents with cat, Read, head, grep, or any other tool.
Reference variables by name only (e.g. $BOOTSTRAP_SERVERS) and verify presence with
test -n "$VAR", never by printing the value..env to .gitignore to avoid accidentally committing credentials.example.com/.example names, fabricated cluster/topic ids. Never substitute real customer
identifiers, hostnames, or credentials when adapting this skill's examples.db79f9e
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.