CtrlK
BlogDocsLog inGet started
Tessl Logo

data-engineering

Data pipeline architecture — ingestion, orchestration, quality, lineage, SLAs. Use when the user asks to 'design data pipelines', 'architect ingestion', 'set up orchestration', 'plan data lake', 'design lakehouse', or mentions Airflow, Dagster, CDC, data lineage, or pipeline SLAs. [EXPLICIT]

SKILL.md
Quality
Evals
Security

Data Engineering: Pipeline Architecture & Data Platform Design

Generic, brand-neutral engineering capability; deep, sourced playbooks live in references/ and knowledge/. [DOC]

Generic, brand-neutral engineering capability; sourced playbooks in references//knowledge/. [DOC]

TL;DR

Data engineering architecture defines how data is ingested, orchestrated, stored, validated, and observed — the backbone infrastructure that feeds analytics, ML, and operational systems. This skill produces data engineering documentation that enables teams to build reliable, scalable, and cost-efficient data platforms [EXPLICIT]

When to Use

  • Designing data ingestion pipelines (batch, streaming, CDC)
  • Architecting pipeline orchestration with dependency management
  • Planning storage architecture (data lake, lakehouse, warehouse zones)
  • Building data quality frameworks with validation and profiling
  • Establishing lineage tracking and pipeline observability
  • Optimizing data platform scalability and cost management

When NOT to Use

  • dbt transformations and data modeling → use analytics-engineering skill
  • Dashboard and reporting architecture → use bi-architecture skill
  • ML model training and serving pipelines → use data-science-architecture skill
  • Application-level software architecture → use software-architecture skill

Sub-capabilities (resource map)

Deep, evidence-tagged playbooks — open the one the task needs (ICM Layer 3, on-demand). [INFERENCE]

Reference
references/full-playbook.md
references/knowledge-graph.mmd
references/pipeline-patterns.md
references/state-of-the-art.md

Procedure

  1. Resolve the sub-capability; open the matching references/ playbook. [EXPLICIT]
  2. Apply its decision tables; pick the strategy explicitly. [EXPLICIT]
  3. Validate against the Quality Criteria and tag every claim. [EXPLICIT]

Quality Criteria

  • Sub-capability resolved to one playbook. [INFERENCE]
  • Claims evidence-tagged. [EXPLICIT]

Contract

  • Aceptación: capability resolved to its reference playbook, applied, validated, evidence-tagged. [EXPLICIT]
  • Límites: · Focuses on data platform and pipelines, not transformation modeling · Does not design consumption layer (dashboards, KPIs) · Does not address ML-specific pipelines (feature. [EXPLICIT]
  • Casos borde: Greenfield Data Platform: Start with managed connectors for quick wins, event-driven architecture for new systems, batch for legacy. Avoid custom connectors until managed optio. [EXPLICIT]
  • Supuestos: · Cloud or on-premise infrastructure is provisioned or being planned · Source systems identified and accessible (or access being negotiated) · Team has Python/SQL skills and famili. [SUPUESTO]
  • Trade-off: Decision Enables Constrains Threshold --- --- --- --- CDC Ingestion Low latency, minimal source impact CDC tool dependency, schema coupling Transactional DB. [EXPLICIT]

Packet

Capas del packet, cargables bajo demanda (disciplina ICM: una capa por vez, nunca todas juntas): references/ guías de profundidad (cargar UNA por etapa) · knowledge/ cuerpo de conocimiento · prompts/ prompts listos · examples/ salida de ejemplo · agents/ subagentes del packet · assets/ recursos estáticos.

Repository
JaviMontano/claude-plugins
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.