CtrlK
BlogDocsLog inGet started
Tessl Logo

metodologia-observability

Observability architecture — logging, tracing, metrics, alerting, SLO/SLI, incident response. Use when the user asks to "design observability", "set up monitoring", "implement tracing", "configure alerting", "define SLOs", "design incident response", or mentions OpenTelemetry, Prometheus, Grafana, ELK, correlation IDs, burn rate, runbooks.

The canonical home for this skill is metodologia-observability in JaviMontano/mao-discovery-framework

SKILL.md
Quality
Evals
Security

Observability Architecture: Instrumentation, Detection & Response

Observability architecture enables teams to understand system behavior from external outputs — logs, traces, and metrics. The skill produces comprehensive observability strategies covering the three pillars, alerting frameworks, and incident response integration that transform raw telemetry into actionable operational intelligence.

Principio Rector

Si no se puede observar, no se puede operar. Si no se puede operar, no existe. Los 3 pilares (logs, metrics, traces) son el mínimo, no el máximo. SLO-based alerting reemplaza threshold alerting, y la respuesta a incidentes empieza con observabilidad — no termina ahí.

Filosofía de Observability

  1. 3 pillars are minimum. Logs, métricas y traces son la base — no el techo. Exemplars conectan métricas con traces. Service maps conectan traces con topología. Sin las tres, hay monitoring — no observability.
  2. SLO-based alerting > threshold alerting. Alertar cuando CPU > 80% es ruido. Alertar cuando el error budget burn rate consume 14.4x en 1 hora es acción. Las alertas sirven al negocio, no a la infraestructura.
  3. Incident response starts with observability. Si el equipo no puede diagnosticar un incidente en <15 minutos con dashboards, traces y logs correlacionados, la arquitectura de observabilidad falló. Post-mortem retroalimenta la instrumentación.

Inputs

The user provides a system or platform name as $ARGUMENTS. Parse $1 as the system/platform name used throughout all output artifacts.

Parameters:

  • {MODO}: piloto-auto (default) | desatendido | supervisado | paso-a-paso
    • piloto-auto: Auto para instrumentation design y log architecture, HITL para SLO targets y alerting thresholds.
    • desatendido: Cero interrupciones. Observability architecture documentada automáticamente. Supuestos documentados.
    • supervisado: Autónomo con checkpoint en collector topology y alert rule design.
    • paso-a-paso: Confirma cada log standard, trace sampling strategy, metric naming, y alert rule.
  • {FORMATO}: markdown (default) | html | dual
  • {VARIANTE}: ejecutiva (~40% — S1 strategy + S4 metrics/dashboards + S5 alerting) | técnica (full 6 sections, default)

Before generating architecture, detect the technology context:

!find . -name "*.yaml" -o -name "*.yml" -o -name "docker-compose*" -o -name "*.tf" -o -name "otel*" | head -20

If reference materials exist, load them:

Read ${CLAUDE_SKILL_DIR}/references/observability-patterns.md

When to Use

  • Designing observability strategy for new systems or platform migrations
  • Implementing structured logging with correlation across services
  • Adding distributed tracing to microservices or event-driven architectures
  • Defining metric collection, naming conventions, and dashboard hierarchies
  • Building SLO-based alerting with burn rate windows
  • Integrating observability with incident response workflows
  • Evaluating or consolidating existing monitoring tooling

When NOT to Use

  • Performance optimization and load testing — use performance-engineering
  • Infrastructure provisioning and platform design — use infrastructure-architecture
  • CI/CD pipeline design and deployment automation — use devsecops-architecture
  • Application code architecture and patterns — use software-architecture

Delivery Structure: 6 Sections

S1: Observability Strategy

Define the overarching approach to system understanding through the three pillars.

Three-pillar assessment: Evaluate current maturity of logging, tracing, and metrics. Score each 1-5 (ad hoc -> optimized).

OpenTelemetry (OTel) Adoption Plan:

  • SDK integration per language (auto-instrumentation for Java, Python, .NET, Node.js)
  • Collector deployment (see topology below)
  • Vendor-agnostic export via OTLP protocol

OTel Collector Topology Decision:

PatternDescriptionWhen to use
Agent (per-node)Sidecar or DaemonSet alongside each app. Lightweight, local processing.Default starting point. Minimizes network hops.
Gateway (centralized)Standalone service receiving from multiple agents. Heavy processing, routing.Tail-based sampling, multi-backend routing, data enrichment.
Hierarchical (Agent + Gateway)Agents handle local batching/filtering; Gateways handle aggregation/sampling.Production recommendation for >10 services.

Critical topology rule for tail-based sampling: Separate Agent from Gateway. All spans from the same trace must reach the same Gateway instance for correct sampling decisions. Use trace-ID-based routing (consistent hashing) at the load balancer in front of Gateways.

Sampling strategies:

  • Head-based (decided at entry): simple, predictable cost, misses interesting traces
  • Tail-based (decided after trace completes): captures slow/error traces, requires Gateway buffering
  • Hybrid: head-based 10% + tail-based for errors and slow traces (recommended)

Retention tiers: Hot 7d (real-time query), warm 30d (aggregated), cold 90d+ (archived, compliance)

Observability-driven development: Instrument first, then code. Define "healthy" in telemetry before writing business logic. New features ship with dashboards and alerts as part of definition of done.

S2: Logging Architecture

Design structured logging with aggregation, correlation, and retention management.

Mandatory structured log fields (JSON):

timestamp, level, service, traceId, spanId, message, environment, version

Log Level Standards:

LevelProduction useDefinition
ERRORAlways onActionable failures requiring intervention
WARNAlways onDegraded but functioning, may need attention
INFOAlways onBusiness events, request lifecycle
DEBUGOff (enable per-service temporarily)Development detail, never in prod by default

Correlation ID propagation: Request ID generated at entry point, passed through all service calls via W3C Trace Context headers. Every log line includes traceId and spanId.

Aggregation pipeline: Collection agent (Vector, Fluentd) -> Processing (filtering, enrichment) -> Storage

Storage backend decision:

  • Elasticsearch: Full-text search, flexible queries. Higher cost at scale.
  • Loki: Label-indexed, cheaper at scale, Grafana-native. No full-text search.
  • CloudWatch/Cloud Logging: Managed, no ops. Limited query power.

Sensitive data: PII masking at collection layer (Vector transforms, Fluentd filters). Never log credentials, tokens, or full credit card numbers.

Retention: ERROR 90d, WARN 30d, INFO 30d, DEBUG 7d (non-prod only)

S3: Distributed Tracing

Implement trace propagation, span design, and cross-signal correlation.

Trace propagation: W3C Trace Context headers across HTTP, gRPC metadata, message queue headers (Kafka record headers, AMQP properties)

Span design: One span per logical operation:

  • HTTP handler (server span)
  • Database query (client span)
  • Cache lookup (client span)
  • Queue publish/consume (producer/consumer span)
  • External API call (client span)

Span attributes: Operation name, status code, error flag, custom business attributes (order ID, tenant ID)

Sampling strategy:

  • 100% for errors (always capture)
  • 100% for traces exceeding p95 latency threshold
  • 1-10% probabilistic for normal success traces
  • Tail-based sampling at Gateway Collector for slow/error traces

Exemplars — Metrics-to-Traces Linking: Exemplars attach a trace/span reference to a specific metric data point. Configure both OTel metric and trace SDKs; record metrics within an active span context. This enables clicking from a latency spike on a dashboard directly to the offending trace — the critical bridge between "what is happening" (metrics) and "why" (traces). Enable exemplars in Prometheus (--enable-feature=exemplar-storage) and Grafana (exemplar data source configuration).

Service map: Auto-generate topology from trace data. Review weekly for unexpected dependencies.

Storage: Tempo (Grafana-native, cost-efficient), Jaeger (mature, Elasticsearch/Cassandra backend), X-Ray (AWS native)

S4: Metrics & Dashboards

Define metric types, naming conventions, collection methods, and dashboard hierarchy.

Metric types: Counters (requests total), Gauges (active connections), Histograms (latency distribution)

Naming convention: <service>_<component>_<metric>_<unit> (e.g., api_http_request_duration_seconds)

Framework methods:

  • RED (for services): Rate, Errors, Duration — every service gets these three
  • USE (for resources): Utilization, Saturation, Errors — for CPU, memory, disk, network

Dashboard hierarchy:

  1. Executive: SLO status across all services (green/yellow/red), error budget remaining
  2. Service: RED metrics per service, dependency health, deployment markers
  3. Component: USE metrics per resource, GC, connection pools, queue depths

Log-Based Metrics: Derive counters and gauges from structured logs without separate instrumentation. Extract ERROR counts per service per minute, parse latency from request logs. Tools: Loki recording rules, Vector transforms, CloudWatch Metric Filters. Useful for legacy systems that only emit logs.

Cardinality Management:

ProblemSolution
URL paths with IDs (/users/123)Normalize to /users/{id}
User/request IDs as labelsRemove; use trace correlation instead
Unbounded enum labelsAllowlist known values, bucket rest as "other"
Per-pod metrics in large clustersAggregate to service level, drill down on demand

Default OTel cardinality limit: 2000 unique time series per metric (configurable via View API). Cardinality explosion is the primary driver of unpredictable observability costs. Enforce per-service observability budgets (e.g., max 5000 active series per service). Use delta temporality for high-cardinality counters.

Dashboard-as-code: Grafana dashboards in JSON/Jsonnet, version-controlled. Deploy with Terraform or grizzly.

S5: Alerting Framework

Build SLO-based alerting with burn rate windows, severity levels, and runbook integration.

Alert Fatigue Prevention — Error Budget Burn Rate Model (Google SRE):

Alert typeBurn rateLong windowShort windowAction
Fast burn14.4x1 hour5 minutesPage immediately (P1)
Medium burn6x6 hours30 minutesPage during hours (P2)
Slow burn1x3 days6 hoursCreate ticket (P3)

This dual-window approach reduces false positives while maintaining sensitivity. Alert only when error budget burn rate is sustained — eliminates transient noise.

Severity Levels:

SeverityCriteriaResponseNotification
P1Customer-visible impact, SLO breach imminentPage immediately, 15min responsePagerDuty/OpsGenie
P2Degraded performance, no SLO breach yetPage during business hoursSlack + on-call
P3Anomaly, potential future issueNext business day ticketEmail + Jira

Alert hygiene:

  • Review noisy alerts monthly. Target: >80% of pages result in action taken.
  • Track alert-to-incident ratio. If <50%, alerts are too noisy.
  • Retire stale alerts quarterly.
  • Suppress during maintenance windows.
  • Aggregate related alerts (group by service + error type).

Runbook linkage: Every alert MUST link to a runbook with: diagnostic steps, likely root causes, remediation actions, escalation path. No alert without a runbook.

S6: Incident Response Integration

Connect observability to incident management with on-call, classification, and post-mortem feedback.

On-call: Rotation schedules (weekly), follow-the-sun for distributed teams, primary + secondary Classification: Severity matrix (impact x urgency), auto-suggest severity from SLO data

Incident timeline: Auto-generated from:

  • Alert timestamps
  • Recent deployments (deploy markers on dashboards)
  • Config changes
  • Related alerts across services

Post-mortem template: Timeline, impact (users affected, duration, error budget consumed), root cause, contributing factors, action items with owners and deadlines

Feedback loops: Post-mortem actions feed into:

  • Alert rule improvements (fewer false positives)
  • Runbook updates (faster resolution next time)
  • Architecture changes (eliminate root cause class)
  • SLO adjustments (if targets were unrealistic)

Blameless culture: Focus on system failures, not individual mistakes. Mandatory post-mortems for P1, optional for P2.


Observability Tool Comparison

CapabilityGrafana Stack (OSS)DatadogNew RelicHoneycomb
MetricsPrometheus/MimirBuilt-inBuilt-inLimited (traces-first)
LogsLokiBuilt-inBuilt-inLimited
TracesTempoBuilt-inBuilt-inCore strength
ExemplarsNativeSupportedSupportedNative
High-cardinality explorationLimitedGoodGoodExcellent
Cost modelInfra + opsPer-host + ingestionPer-GB ingestedPer-event
Vendor lock-inNone (OTel native)MediumMediumLow (OTel)
Best forPlatform teams, cost-conscious, full controlFull-stack teams, fast setupLegacy + modern mixedDebugging complex distributed systems

Selection criteria: <20 services and no platform team -> managed (Datadog/New Relic). Platform engineering team present -> Grafana stack. Complex distributed debugging priority -> Honeycomb. Always use OTel SDK regardless of backend for portability.


Trade-off Matrix

DecisionEnablesConstrainsWhen to Use
Full trace samplingComplete visibilityHigh storage cost, processing overheadDebugging, low-traffic, compliance
Tail-based samplingCaptures slow/error reliablyCollector buffering, added latencyHigh-traffic, anomaly focus
Managed platformFast setup, unified UIVendor lock-in, cost at scaleSmall-medium teams, rapid value
Open-source stackNo lock-in, full controlOps burden, integration workLarge teams with platform capacity
SLO-based alertingReduced noise, business-alignedRequires SLO definition disciplineMature teams with defined service levels

Assumptions

  • Services can be instrumented (source access or sidecar injection available)
  • Network allows telemetry collection (firewall rules, egress for cloud collectors)
  • Team has on-call rotation or is willing to establish one
  • Budget exists for observability tooling (storage, compute, licensing)

Limits

  • Does not design application architecture
  • Does not perform load testing or capacity planning
  • Does not implement security monitoring or SIEM
  • Observability data quality depends on instrumentation discipline across all teams
  • Cost estimates are directional; actual costs depend on data volumes and retention

Edge Cases

Greenfield System: Instrument from day one. Embed OTel SDK in service templates. Define logging and tracing standards before first deployment.

Legacy Monolith: Start with infrastructure metrics and access logs. Add structured logging incrementally. Use APM agents for automatic instrumentation. Trace boundaries at external calls.

Serverless / FaaS: Cold starts complicate tracing. Use OTel Lambda layers or vendor-native tracing (X-Ray, Cloud Trace). Push metrics (no scrape endpoint). Log to stdout with structured format.

Multi-Cloud or Hybrid: Normalize telemetry with OTel. Centralized Gateway Collector aggregating across clouds. Standardize naming conventions regardless of provider.

High-Cardinality Environments: Microservices with many endpoints and tenants. Use label allowlists, drop unused labels at collection, enforce per-service series budgets.


Validation Gate

Before finalizing delivery, verify:

  • All three pillars covered with specific tooling and OTel Collector topology defined
  • Exemplars configured linking metrics to traces
  • Log levels have clear definitions with production defaults
  • Trace sampling strategy balances cost and visibility (hybrid recommended)
  • Metric naming follows consistent conventions with cardinality limits enforced
  • Dashboard hierarchy exists (executive SLO, service RED, component USE)
  • Alerts use SLO-based burn rate with multi-window configuration
  • Alert-to-runbook linkage is mandatory (no alert without runbook)
  • Alert fatigue metrics tracked (>80% pages result in action)
  • Incident response connects alerts -> on-call -> post-mortem -> improvement loop

Knowledge Graph

graph TD
    subgraph Core
        OBS[Observability Architecture]
    end

    subgraph Inputs
        I1[System Topology] --> OBS
        I2[Existing Monitoring Tools] --> OBS
        I3[SLO Requirements] --> OBS
        I4[Incident History] --> OBS
    end

    subgraph Outputs
        OBS --> O1[Three-Pillar Strategy]
        OBS --> O2[OTel Collector Topology]
        OBS --> O3[Logging Standards]
        OBS --> O4[Tracing Design]
        OBS --> O5[Alerting Framework]
        OBS --> O6[Incident Response Integration]
    end

    subgraph Related Skills
        RS1[performance-engineering] -.-> OBS
        RS2[infrastructure-architecture] -.-> OBS
        RS3[devsecops-architecture] -.-> OBS
        RS4[software-architecture] -.-> OBS
    end

Output Templates

Formato MD (default):

# Observability Architecture: {system_name}
## S1: Observability Strategy
### Maturity Assessment | OTel Adoption Plan | Sampling | Retention

## S2: Logging Architecture
### Structured Fields | Level Standards | Aggregation | Storage

## S3: Distributed Tracing
### Propagation | Span Design | Exemplars | Service Map

## S4: Metrics & Dashboards
### RED/USE | Naming Conventions | Dashboard Hierarchy | Cardinality

## S5: Alerting Framework
### Burn Rate Model | Severity Levels | Runbook Linkage

## S6: Incident Response Integration
### On-Call | Classification | Post-Mortem | Feedback Loops

Formato XLSX: Matriz de controles de observabilidad: filas por servicio/componente, columnas por pilar (logs, traces, metrics), con estado de implementacion, herramienta asignada, y gaps identificados. Incluye hoja de SLOs con burn rate calculations.

Formato DOCX (bajo demanda):

  • Filename: {fase}_Observability_Architecture_{cliente}_{WIP}.docx
  • Generado via python-docx con MetodologIA Design System v5. Portada con logo y metadatos, TOC automatico, headers/footers con nombre del skill y numeracion, tablas zebra, titulos Poppins navy, cuerpo Montserrat, acentos gold.

Formato PPTX (bajo demanda):

  • Filename: {fase}_Observability_Architecture_{cliente}_{WIP}.pptx
  • Generado via python-pptx con MetodologIA Design System v5. Slide master navy gradient, titulos Poppins, cuerpo Montserrat, acentos gold. Max 20 slides variante ejecutiva / 30 variante tecnica. Speaker notes con referencias de evidencia [DOC]/[INFERENCIA]/[SUPUESTO].

Evaluacion

DimensionPesoCriterio (7/10 minimo)
Trigger Accuracy10%Se activa ante keywords de observability, monitoring, tracing, alerting; no se confunde con performance testing
Completeness25%Los 3 pilares cubiertos con tooling concreto, OTel topology definida, alerting con burn rate multi-window
Clarity20%Collector topology, log level standards, y severity levels son inequivocos y operacionalizables
Robustness20%Edge cases (greenfield, legacy, serverless, multi-cloud, high-cardinality) tienen estrategia especifica
Efficiency10%Variante ejecutiva entrega estrategia + dashboards + alerting en ~40% del contenido
Value Density15%Cada seccion produce configuracion aplicable: OTel config, alert rules, dashboard JSON, runbook templates

Umbral minimo: 7/10 en cada dimension. Composite ponderado >= 7.0 para considerar el output aceptable.


Output Format Protocol

FormatDefaultDescription
markdownYesRich Markdown + Mermaid diagrams. Token-efficient.
htmlOn demandBranded HTML (Design System). Visual impact.
dualOn demandBoth formats.

Default output is Markdown with embedded Mermaid diagrams. HTML generation requires explicit {FORMATO}=html parameter.

Output Artifact

Primary: A-01_Observability_Architecture.html — Executive summary, three-pillar strategy, OTel Collector topology, logging standards, tracing design, metric taxonomy, alerting framework, incident response integration.

| HTML | {fase}_Observability_Architecture_{cliente}_{WIP}.html | Mismo contenido en HTML branded (Design System MetodologIA v5). Self-contained, WCAG AA, responsive. Tipo: Light-First Technical. Incluye dashboard de madurez por pilar, topología de collector interactiva, y tabla de burn rate alerts. |

Secondary: OTel Collector configuration (agent + gateway), Grafana dashboard JSON, burn rate alert rules, runbook templates, post-mortem template.


Autor: Javier Montaño | Última actualización: 12 de marzo de 2026

Repository
JaviMontano/mao-pm-apex
Last updated
First committed

Canonical home

JaviMontano/mao-discovery-framework
In sync

since Aug 28, 2026

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.