CtrlK
BlogDocsLog inGet started
Tessl Logo

chaos-experiment

Create and manage chaos experiments using Harness Chaos Engineering via MCP. Run resilience tests like pod deletion, CPU stress, and network faults. Use when user says "chaos experiment", "chaos engineering", "resilience test", "chaos test", or wants to test system reliability.

63

Quality

74%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/claude/skills/chaos-experiment/SKILL.md
SKILL.md
Quality
Evals
Security

Chaos Experiment

Create and manage chaos experiments using Harness Chaos Engineering via MCP.

Instructions

Step 1: Check Infrastructure

Call MCP tool: harness_list
Parameters:
  resource_type: "chaos_infrastructure"
  org_id: "<organization>"
  project_id: "<project>"

Step 2: Browse Templates

Call MCP tool: harness_list
Parameters:
  resource_type: "chaos_experiment_template"
  org_id: "<organization>"
  project_id: "<project>"

Step 3: List Existing Experiments

Call MCP tool: harness_list
Parameters:
  resource_type: "chaos_experiment"
  org_id: "<organization>"
  project_id: "<project>"

Step 4: Create Experiment

Call MCP tool: harness_create
Parameters:
  resource_type: "chaos_experiment"
  org_id: "<organization>"
  project_id: "<project>"
  body: <experiment definition>

Step 5: Run Experiment

Call MCP tool: harness_execute
Parameters:
  resource_type: "chaos_experiment"
  action: "run"
  resource_id: "<experiment_id>"
  org_id: "<organization>"
  project_id: "<project>"

Step 6: Monitor Results

Call MCP tool: harness_list
Parameters:
  resource_type: "chaos_experiment_run"
  org_id: "<organization>"
  project_id: "<project>"

Get specific run details:

Call MCP tool: harness_get
Parameters:
  resource_type: "chaos_experiment_run"
  resource_id: "<run_id>"

Step 7: Check Probes

Call MCP tool: harness_list
Parameters:
  resource_type: "chaos_probe"
  org_id: "<organization>"
  project_id: "<project>"

Common Experiment Types

  • Pod Delete - Kill pods to test recovery
  • Pod CPU Hog - Stress CPU to test throttling
  • Pod Memory Hog - Consume memory to test OOM handling
  • Pod Network Loss - Simulate network failures
  • Pod Network Latency - Add artificial latency
  • Node Drain - Drain K8s nodes
  • EC2 Stop - Stop AWS EC2 instances
  • ECS Task Stop - Stop ECS tasks

Chaos Resource Types

Resource TypeOperationsDescription
chaos_experimentlist, get, create, update, delete, runExperiments
chaos_experiment_runlist, getRun history/results
chaos_experiment_templatelist, getPre-built templates
chaos_infrastructurelist, getTarget infrastructure
chaos_probelist, getHealth probes

Examples

  • "Show me all chaos experiments" - List chaos_experiment
  • "Create a pod-delete experiment for checkout-service" - Create chaos_experiment
  • "Run the weekly resilience test" - Execute run action
  • "What were the results of the last chaos run?" - Get chaos_experiment_run

Performance Notes

  • Review existing experiments before creating duplicates. Check for similar fault types targeting the same service.
  • Wait for experiment completion before analyzing results. Do not draw conclusions from partial runs.
  • Verify the target infrastructure and service are healthy before running chaos experiments.

Troubleshooting

Experiment Won't Run

  • Verify chaos infrastructure is connected and active
  • Check target application/namespace exists
  • Ensure RBAC permissions for chaos operations

Probes Failing

  • Check probe endpoints are accessible
  • Verify probe timeout settings
  • Review probe type matches expected behavior
Repository
harness/harness-ai
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.