CtrlK
BlogDocsLog inGet started
Tessl Logo

observability-operations

Production observability and operations: metrics/dashboards and SLO/SLI, alerting design, structured logging architecture, health-check probes, incident response and post-mortems. [EXPLICIT] Trigger: 'observability', 'monitoring', 'alerting', 'logging', 'slo', 'sli', 'incident response', 'post-mortem'.

SKILL.md
Quality
Evals
Security

Observability & Operations

"You cannot operate what you cannot see; you cannot alert on what you cannot measure." [INFERENCE]

TL;DR

Designs production observability and operations: metrics with SLO/SLI and dashboards, alerting that avoids fatigue, structured logging with privacy masking, health-check probes, and incident response with blameless post-mortems. Turns a deployed system into an operable one. [EXPLICIT]

When to use

  • Defining SLOs/SLIs and dashboards for a service. [EXPLICIT]
  • Designing alerting rules that page on symptoms, not noise. [EXPLICIT]
  • Architecting structured logging (aggregation, retention, PII masking, sampling). [EXPLICIT]
  • Adding health-check liveness/readiness probes and synthetic monitors. [EXPLICIT]
  • Writing an incident-response + post-mortem playbook. [EXPLICIT]

Sub-capabilities (resource map)

CapabilityReference
Metrics + dashboards + SLOreferences/monitoring-setup.md
Alertingreferences/alerting-strategy.md
Loggingreferences/log-management.md
Health checksreferences/health-check-automation.md
Incident responsereferences/incident-response.md

Procedure

  1. Define SLIs (latency, error rate, saturation) and SLO targets first. [DOC]
  2. Alert on SLO burn / symptoms, not every metric; set severity + escalation. [DOC]
  3. Log structured with correlation IDs; mask PII; set retention + sampling. [DOC]
  4. Probe health (liveness vs readiness) and add synthetic checks for critical flows. [DOC]
  5. Rehearse incidents: classification, roles, comms, blameless post-mortem with actions. [EXPLICIT]

Quality Criteria

  • SLOs/SLIs defined before dashboards. [DOC]
  • Alerts are symptom/burn-based with severity + escalation (low false-page rate). [INFERENCE]
  • Logs structured, correlated, PII-masked, retained per policy. [DOC]
  • Liveness ≠ readiness; critical flows have synthetic monitors. [DOC]
  • Claims evidence-tagged. [EXPLICIT]

Anti-Patterns

  • Alerting on every metric → alert fatigue → ignored pages. [INFERENCE]
  • Logging PII in plaintext. [DOC]
  • Conflating liveness and readiness (restarts a healthy-but-busy pod). [DOC]

Contract

  • Aceptación: SLO/SLI antes de dashboards; alertas por síntoma/burn con escalación; logs estructurados+enmascarados; probes correctas; incident playbook con post-mortem accionable. [EXPLICIT]
  • Límites: observabilidad y operación; no es perf-tuning de código (eso es otra skill) ni CI/CD (cicd-release-engineering). [EXPLICIT]
  • Casos borde: alta cardinalidad de métricas → costo/explosión; logs con PII → riesgo de cumplimiento. [INFERENCE]
  • Supuestos: acceso a la plataforma de métricas/logs del usuario. [SUPUESTO]
  • Trade-off: más señal cuesta almacenamiento/ruido; el arte es alta cobertura con bajo false-page. [EXPLICIT]

Related Skills

  • cicd-release-engineering — deploy events feed dashboards/alerts. [EXPLICIT]
  • web-infra-foundation — infra it monitors (DNS/SSL/CDN health). [EXPLICIT]
  • application-security — audit logging overlaps with security audit. [EXPLICIT]

Packet

Capas del packet, cargables bajo demanda (disciplina ICM: una capa por vez, nunca todas juntas): references/ guías de profundidad (cargar UNA por etapa) · knowledge/ cuerpo de conocimiento · prompts/ prompts listos · examples/ salida de ejemplo · agents/ subagentes del packet · assets/ recursos estáticos.

Repository
JaviMontano/claude-plugins
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.