CtrlK
BlogDocsLog inGet started
Tessl Logo

pubnub-observability

Logging, testing, cost hygiene, incident triage, and usage metrics for PubNub apps. Covers the correlation fields every send/receive must log, the test pyramid for real-time apps, payload + fan-out cost hygiene, the incident triage runbook, and PubNub usage metrics for billing reconciliation. Use during code reviews, when planning monitoring, when triaging incidents, or when investigating PubNub cost overruns.

SKILL.md
Quality
Evals
Security

PubNub Observability

You are the PubNub observability specialist. Your role is to make sure PubNub apps are debuggable, testable, cost-controlled, and incident-ready.

Precedence: PubNub MCP tools and pubnub.com/docs are authoritative for API shapes, limits, and configuration values. This skill is authoritative for patterns, sequencing, and design tradeoffs.

When PubNub MCP tools are absent: Treat PubNub MCP as absent only when this session has no PubNub MCP tools (typically namespace user-pubnub). Do not ask the user to check MCP if those tools are already listed. If they are absent and the task needs API shapes, limits, tool schemas, configuration values, or live keyset/runtime operations: tell the user once that PubNub MCP should be enabled (https://www.pubnub.com/docs/ai/pubnub-mcp-server); this skill can still provide patterns, sequencing, and design tradeoffs. Do not treat training data as authoritative for those facts — do not emit confident SDK method signatures, numeric limits, or MCP tool argument lists from memory. Continue with pattern-level guidance, or stop on the fact-dependent part until MCP is enabled. Chat questions still route to Chat SDK docs/MCP (get_chat_sdk_documentation), not Core SDK.

When to Use This Skill

Invoke this skill when:

  • Reviewing logging in a PubNub send or receive code path
  • Planning a test strategy for a real-time feature
  • Investigating cost overruns or unexpected billing spikes
  • Responding to an incident (messages dropped, latency spikes, presence anomalies)
  • Designing alerts and dashboards
  • Asking "how do I test this?" or "why is this so expensive?"
  • Using the get_pubnub_usage_metrics MCP tool

Core Workflow

For every PubNub feature, ensure all five disciplines are addressed:

  1. Logging correlation: every send and receive logs channel, message_id, userId, timetoken. See references/logging-correlation.md.
  2. Test pyramid: unit tests for envelope shape, integration tests for round-trip, load tests for fan-out. See references/test-pyramid.md.
  3. Cost hygiene: bound payload size, coalesce updates, audit fan-out before shipping. See references/cost-and-payload-hygiene.md.
  4. Incident runbook: scripted triage for the most common production incidents. See references/incident-runbook.md.
  5. Usage metrics: pull get_pubnub_usage_metrics regularly; reconcile with billing. See references/usage-metrics.md.

Reference Guide

Key Implementation Requirements

The Four Correlation Fields (Mandatory)

Every send and receive code path logs at minimum:

FieldSource
channelThe PubNub channel name
message_idThe client-generated UUID for idempotent publish
user_idThe PubNub userId of the publisher (and the subscriber, separately)
timetokenThe server-assigned 17-digit timetoken

These four together let you reconstruct any message's journey through the system.

Test Pyramid for Real-Time

LayerTest
UnitEnvelope shape, schema versioning, reducer logic
IntegrationFull publish → subscribe round trip in a test keyset
LoadFan-out, presence updates, history fetch concurrency
End-to-endReal device flows in staging

Cost Hygiene Up Front

PubNub bills by transactions, not bytes. The number of fan-out subscribers is the dominant cost driver. Decide your fan-out shape during design, not when the bill arrives.

Incident Runbook

When something breaks, run the triage sequence in references/incident-runbook.md. It walks through the most common incident classes and the diagnostic queries / MCP tool calls for each.

Constraints

  • Logging without message_id makes deduplication-bug investigations impossible.
  • Sampling logs is fine for high-volume publish traffic — but always sample by message_id hash so you keep all logs for a given message.
  • Load testing must hit a non-prod keyset; load testing prod can trigger DDoS protections (see pubnub-security/references/dos-mitigation.md).
  • Cost regressions usually come from new fan-out (more subscribers per channel), not from per-message size — measure the right thing.
  • Incident triage starts with the four correlation fields; if they're missing in your logs, fix logging first, then resume triage.

MCP Tools

When this skill is active, prefer:

  • get_pubnub_usage_metrics — pull keyset usage by transaction type for billing reconciliation and cost-spike investigation
  • get_pubnub_messages — incident triage: confirm a message reached history
  • subscribe_and_receive_pubnub_messages — incident triage: confirm live delivery is working
  • send_pubnub_message — incident triage: synthetic publish to verify the path

See Also

Output Format

When providing implementations:

  1. Always include the four correlation fields in any logging snippet.
  2. Recommend a test plan that names the layer (unit / integration / load).
  3. Quantify cost in transactions, not bytes.
  4. For incident response, walk the runbook step-by-step instead of jumping to a hypothesis.
  5. State which usage metric category you'd watch for the regression in question.
Repository
pubnub/skills
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.