CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/guardrail-metrics-reference

Pure-reference catalog of guardrail-metric methodology for online controlled experiments. Defines guardrail metrics (metrics that must NOT degrade for an experiment to ship, even if the primary metric improves), the standard guardrail set (latency / errors / engagement / opt-out), the relationship to OEC (Overall Evaluation Criterion) per Kohavi et al., and pre-commitment of the metric set. The quantitative evaluation mechanics (per-metric alert/block thresholds, Bonferroni / Benjamini-Hochberg multiple-comparison correction) live in references/. Use when designing the metric set for a new experiment, auditing existing experiment configs, or reviewing experiment results before ship-decisions.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured reference catalog with actionable config examples, a clear sequenced workflow, and clean one-level-deep progressive disclosure. Its only weakness is conciseness: a few decorative passages (the Microsoft quote, ISBN, repeated definitions) could be trimmed.

Suggestions

Cut the Microsoft Experimentation Platform quote and SRM tangent — they are decorative and the SRM concept belongs to the sibling peeking-problem-reference, not this skill.

State the guardrail/OEC definition once (Overview) and reference it rather than re-explaining it in "The OEC vs guardrail relationship"; drop the inline ISBN (keep it only in References).

Tighten the Limitations bullets to one line each so the section reads as a checklist rather than prose.

DimensionReasoningScore

Conciseness

Mostly efficient (dense tables, YAML example) but includes decorative or redundant context — the Microsoft/SRM quote block, repeated ISBN, and a second restatement of the guardrail definition — that Claude does not need; not yet at the "every token earns its place" anchor.

2 / 3

Actionability

Provides concrete, copy-paste-ready guidance: a worked YAML experiment config, a canonical guardrail-set table with directions, and an anti-pattern→fix table; per the instruction-only scoring note, absence of executable code is not penalized when guidance is this actionable.

3 / 3

Workflow Clarity

The 5-step "How to use this reference" gives a clear sequence (pick set → declare → set thresholds → correct alpha → gate ship) with an explicit ship-gate checkpoint in step 5; not a destructive/batch operation so heavy feedback loops are not required.

3 / 3

Progressive Disclosure

Overview lives in SKILL.md while the deep quantitative mechanics (thresholds, Bonferroni/BH) are split into one clearly-signaled, one-level-deep reference (references/thresholds-and-corrections.md) that exists as a real file.

3 / 3

Total

11

/

12

Passed

Description

100%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that states concrete capabilities, includes an explicit Use-when trigger with natural task phrases, and occupies a distinctive niche. It hits the top anchor on every dimension with no verbosity or over-claiming.

DimensionReasoningScore

Specificity

Lists multiple concrete capabilities: "Defines guardrail metrics", "the standard guardrail set (latency / errors / engagement / opt-out)", "relationship to OEC", and "pre-commitment of the metric set" — matching the multiple-specific-actions anchor.

3 / 3

Completeness

Explicitly answers both what ("Pure-reference catalog of guardrail-metric methodology ... Defines guardrail metrics ...") and when ("Use when designing the metric set ..."), with an explicit Use-when trigger.

3 / 3

Trigger Term Quality

The clause "Use when designing the metric set for a new experiment, auditing existing experiment configs, or reviewing experiment results before ship-decisions" covers natural phrases a user would say, the domain's common task entry points.

3 / 3

Distinctiveness Conflict Risk

Occupies a clear niche — guardrail metrics for online controlled experiments — with triggers scoped to experiment design/audit/review that are unlikely to fire for unrelated skills.

3 / 3

Total

12

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents