CtrlK
BlogDocsLog inGet started
Tessl Logo

ml-failure-audit

General workflow for auditing ML CI failures, experiment regressions, training run failures, golden metric failures, and telemetry-backed ML work-product claims from local repositories, logs, metrics, configs, and artifacts. Use when Codex needs to decide whether an ML failure is a model/convergence issue, correctness bug, data/config issue, infrastructure/runtime issue, evaluation/gating policy issue, or unsupported claim, and produce structured evidence-backed outputs.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill body is a well-structured, lean overview with a clearly sequenced workflow, an executable helper script, and clean one-level-deep references to real bundle files. Minor gains are available by tightening a few redundant framing sentences and surfacing the validate-and-retry loop inline.

Suggestions

Add an explicit inline validation checkpoint (e.g., 'Validate the output: confirm JSON parses and cited formulas match numbers before finalizing') rather than relying on references/output_guidance.md for the validate-and-retry loop.

Trim slightly redundant framing such as 'The script is intentionally generic. It extracts...' to keep the helper-script section as tight as the rest of the body.

Optionally include a one-line concrete example of a recompute formula (e.g., relative error or throughput) inline so the 'Recompute key facts' step is directly executable.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence, with only minor over-explanation (e.g., 'The script is intentionally generic. It extracts...') that could be trimmed, fitting the efficient-with-minor-padding anchor.

4 / 5

Actionability

Provides a concrete, executable, real script invocation with real flags and a complete script, plus concrete workflow directives (cite exact files, record formulas), with minor gaps since the workflow steps are instructional rather than copy-paste code.

4 / 5

Workflow Clarity

Six steps are clearly sequenced (locate, classify, recompute, trace, decide, write) with recomputation and verification guidance present, but the explicit validate-then-retry feedback loop lives in the referenced file rather than as an inline checkpoint.

4 / 5

Progressive Disclosure

The body is a clear overview with well-signaled, one-level-deep references to real files (references/workflow.md, references/output_guidance.md, scripts/collect_failure_evidence.py), with content appropriately split across them.

5 / 5

Total

17

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, comprehensive, and clearly answers both what the skill does and when to use it with concrete trigger language. Keyword coverage is strong but a few natural synonyms and tooling terms are missing.

DimensionReasoningScore

Specificity

Enumerates concrete auditing actions and a comprehensive set of specific failure types ('ML CI failures, experiment regressions, training run failures, golden metric failures') plus concrete input sources (repos, logs, metrics, configs, artifacts), matching the comprehensive-coverage anchor.

5 / 5

Completeness

Explicitly states both what it does (audit ML failures, classify, produce structured evidence-backed outputs) and when to use it with a concrete 'Use when Codex needs to decide whether an ML failure is...' trigger clause.

5 / 5

Trigger Term Quality

Contains natural user-facing keywords ('ML CI failures', 'training run failures', 'golden metric', 'logs', 'metrics') but omits some common synonyms and tooling terms (e.g., W&B/MLflow/TensorBoard, 'training loss', specific extensions), so coverage is good rather than comprehensive.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear ML-failure-audit niche with a specific failure-type taxonomy and telemetry-backed-claim framing, leaving only minor overlap risk with generic CI-debugging or code-review skills.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
eigent-ai/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.