CtrlK
BlogDocsLog inGet started
Tessl Logo

ai-testing-strategy

Comprehensive testing strategy for AI systems — testing scope matrix (6 types x 6 layers), model prediction testing, data quality testing, compliance and fairness testing, integration approaches, and CI/CD test automation. This skill should be used when the user asks to 'define AI testing strategy', 'test ML models', 'design data quality tests', 'plan fairness testing', 'test AI pipelines', 'design integration tests for ML', or mentions adversarial testing, drift simulation, model regression testing, bias testing, explainability testing, or AI test automation. [EXPLICIT]

SKILL.md
Quality
Evals
Security

AI Testing Strategy: Comprehensive Verification for AI-Enabled Systems

Generic, brand-neutral engineering capability; deep, sourced playbooks live in references/ and knowledge/. [DOC]

Generic, brand-neutral engineering capability; sourced playbooks in references//knowledge/. [DOC]

TL;DR

AI testing strategy defines how to verify that an AI system behaves correctly, fairly, securely, and reliably across all layers — from data ingestion through model inference to production monitoring. This skill produces a testing strategy document covering the testing scope matrix, model and prediction tests, data quality tests, compliance and fairness tests, integration approaches, and CI/CD test automation for AI pipelines [EXPLICIT]

When to Use

  • Defining a comprehensive testing strategy for new or existing AI systems
  • Designing model validation tests (accuracy, fairness, robustness, explainability)
  • Planning data quality tests for AI pipelines (schema, distribution, lineage)
  • Implementing compliance and fairness testing (bias detection, audit trails, governance)
  • Selecting integration testing approaches for AI systems (top-down, bottom-up, parallel, harness)
  • Automating AI tests within CI/CD pipelines
  • Evaluating test coverage gaps in existing AI systems

When NOT to Use

  • Internal module structure and layer architecture -> ai-software-architecture
  • CONOPS and operational concept -> ai-conops
  • Pipeline design and CI/CD deployment strategy -> ai-pipeline-architecture
  • Design pattern selection -> ai-design-patterns
  • GenAI/LLM-specific testing (hallucination, RAG quality) -> genai-architecture
  • Traditional software testing without AI context -> testing-strategy

Sub-capabilities (resource map)

Deep, evidence-tagged playbooks — open the one the task needs (ICM Layer 3, on-demand). [INFERENCE]

Reference
references/ai-test-types.md
references/full-playbook.md
references/integration-approaches.md
references/testing-matrix.md

Procedure

  1. Resolve the sub-capability; open the matching references/ playbook. [EXPLICIT]
  2. Apply its decision tables; pick the strategy explicitly. [EXPLICIT]
  3. Validate against the Quality Criteria and tag every claim. [EXPLICIT]

Quality Criteria

  • Sub-capability resolved to one playbook. [INFERENCE]
  • Claims evidence-tagged. [EXPLICIT]

Contract

  • Aceptación: capability resolved to its reference playbook, applied, validated, evidence-tagged. [EXPLICIT]
  • Límites: · Focuses on testing strategy, not test implementation code (see testing frameworks documentation) · Does not design pipeline architecture (see ai-pipeline-architecture) -. [EXPLICIT]
  • Casos borde: No Ground Truth Available: Some AI systems (unsupervised, generative) lack clear ground truth. Use proxy metrics (human evaluation, downstream task performance), A/B testing ag. [EXPLICIT]
  • Supuestos: · AI system has defined requirements with measurable thresholds (AP, NF, SEC, CP metrics) · Test infrastructure (compute, storage) budget is allocated · Team has access to represen. [SUPUESTO]
  • Trade-off: Decision Enables Constrains When to Use --- --- --- --- Full matrix coverage Comprehensive quality assurance High test maintenance cost, slow pipeline Regul. [EXPLICIT]

Packet

Capas del packet, cargables bajo demanda (disciplina ICM: una capa por vez, nunca todas juntas): references/ guías de profundidad (cargar UNA por etapa) · knowledge/ cuerpo de conocimiento · prompts/ prompts listos · examples/ salida de ejemplo · agents/ subagentes del packet · assets/ recursos estáticos.

Repository
JaviMontano/claude-plugins
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.