CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/steady-state-hypothesis-validator

Validates a chaos experiment's steady-state hypothesis before execution: checks that each probe metric is measurable and observable, that a recent baseline exists, that tolerances are numerically meaningful and SLI-backed, that the measurement window is defined, and that the chosen metrics would actually move under the target failure mode. Use when a chaos experiment has been authored (via chaos-experiment-author) and the team needs a pre-flight verdict before running the drill in any environment.

70

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Overview
Quality
Evals
Security
Files

Quality

Content

72%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality, actionable body with concrete YAML, a worked example, and a precise output template. It loses a little on conciseness for a duplicated block-quote/schema restatement and on workflow clarity for lacking an explicit revise-and-re-run feedback loop.

Suggestions

Trim the verbatim principlesofchaos.org Principle 1 block-quote to a one-line paraphrase and drop the in-body restatement of tolerance-gate mechanics, since both are covered in the reference file.

Add an explicit feedback loop to the output/workflow section — e.g. after a FAIL verdict, revise the flagged probe and re-run the pre-flight checks until SOUND — to lift workflow clarity.

Consider moving the hard-reject table's per-check mapping into the reference file or a dedicated section to keep the main flow lean.

DimensionReasoningScore

Conciseness

Mostly efficient and well-structured, but the verbatim principlesofchaos.org Principle 1 block-quote and the in-body restatement of tolerance-gate mechanics (already detailed in the reference file) are mild redundancy that could be tightened.

2 / 3

Actionability

Provides copy-paste-ready probe YAML, concrete pass/fail signals per check, specific diagnostic questions, and an exact output-format template with a fully worked example.

3 / 3

Workflow Clarity

The five checks are clearly sequenced with a hard-reject validation gate and per-probe verdict rows, but no explicit fix→re-validate feedback loop is formalized, which the feedback-loops guidance keeps below a clean 3.

2 / 3

Progressive Disclosure

Clear overview with a well-signaled one-level-deep reference (references/chaostoolkit-tolerance.md, which exists and holds the promised schema/tolerance detail); content appropriately split with easy navigation.

3 / 3

Total

10

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, well-scoped description: it names concrete actions, gives natural trigger terms, includes an explicit 'Use when' clause, and carves out a distinct niche from its sibling authoring skill. No vague fluff or over-claims.

DimensionReasoningScore

Specificity

Enumerates five distinct concrete checks (measurable/observable, baseline exists, SLI-backed tolerance, measurement window, metric moves under failure), matching the 'lists multiple specific concrete actions' anchor.

3 / 3

Completeness

Explicitly answers what (the five pre-flight checks) and when via a clear 'Use when a chaos experiment has been authored... and the team needs a pre-flight verdict' clause.

3 / 3

Trigger Term Quality

Uses natural domain terms users would say — 'chaos experiment', 'steady-state hypothesis', 'pre-flight verdict', 'running the drill' — with good coverage and no jargon-only phrasing.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear niche (steady-state hypothesis pre-flight validation) and disambiguates from the sibling chaos-experiment-author skill, making wrong-skill triggering unlikely.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents