CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/chaos-results-reporter

Aggregates chaos drill verdicts over time into a resilience trend report - per-experiment hypothesis-held / blast-radius / time-to-detect / time-to-recover, degradation trends across runs, action items, and a stakeholder summary. Use when a team has completed one or more chaos drills and needs a structured trend report showing whether resilience is improving, degrading, or stable across iterations.

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Overview
Quality
Evals
Security
Files

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with a well-sequenced, validated workflow and clean one-level progressive disclosure. Its only weakness is mild verbosity from explaining chaos-engineering and ISO-recoverability concepts Claude already knows.

Suggestions

Trim the opening Chaos Engineering definition blockquote and the ISTQB/ISO 25010 recoverability mapping to a one-line citation; Claude already knows these concepts and only needs the principle names used in the workflow.

Consider collapsing the two 'Differentiation axis' paragraphs into a single two-bullet contrast to save tokens while preserving the boundary against chaos-experiment-author and live drill runs.

DimensionReasoningScore

Conciseness

The body is mostly efficient, but the opening blockquote defining Chaos Engineering and the ISTQB/ISO recoverability mapping explain concepts Claude already knows, so it could be tightened; it is not lean enough for the top anchor.

2 / 3

Actionability

Concrete executable guidance throughout - explicit formulas ('held_rate = count of PASSED runs / total runs'), specific thresholds (20% above/below, 10 percentage points, 3 consecutive PASSED runs), exact output path, and a fully worked example - copy-paste ready decision logic.

3 / 3

Workflow Clarity

A clear 7-step sequence with an upfront 'How to use' overview, explicit validation checkpoints (hard-reject rule for zero reports, INCOMPLETE flagging for missing fields), and an anti-patterns table providing error-recovery guidance.

3 / 3

Progressive Disclosure

Overview in SKILL.md with a single one-level-deep reference (references/drill-fields-and-signals.md) that exists on disk and is clearly signaled by inline links in Steps 1 and 4; field-extraction and signal catalog are appropriately split out.

3 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concrete, trigger-rich, third-person, and well-differentiated, explicitly covering both what the skill does and when to invoke it. It hits the top anchor on every dimension with no fluff or over-claims.

DimensionReasoningScore

Specificity

Lists multiple concrete actions: 'per-experiment hypothesis-held / blast-radius / time-to-detect / time-to-recover, degradation trends across runs, action items, and a stakeholder summary', matching the anchor that lists several specific concrete actions.

3 / 3

Completeness

Clearly answers both what (aggregates drill verdicts into a trend report) and when with an explicit 'Use when a team has completed one or more chaos drills and needs a structured trend report' trigger, matching the top anchor.

3 / 3

Trigger Term Quality

Natural terms a user would say are well covered - 'chaos drills', 'trend report', 'resilience', 'improving, degrading, or stable' - rather than technical jargon, satisfying the anchor for good coverage of natural terms.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear niche (post-hoc multi-drill trend aggregation) with distinct triggers; the body even differentiates it from chaos-experiment-author and live drill runs, so it is unlikely to fire for the wrong skill.

3 / 3

Total

12

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Reviewed

Table of Contents