CtrlK
BlogDocsLog inGet started
Tessl Logo

triage-agent-eval-failures

Triage failing @novu/agent-evals scenarios to decide whether a failure is real or flaky, and whether to fix the playbook/prompt or the test (grader, tape, scenario, or judge). Use when an agent-evals scenario fails, when the user asks why an eval is red, or when deciding whether to fix the test or the prompt.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable triage workflow: concrete commands, exact file locations, a decision table, an explicit output format, and validation checkpoints with re-run feedback loops. The two real defects are the missing referenced file (reference.md) that dead-ends the only external pointer, and a duplicated command block in Rule 0.

Suggestions

Ship the referenced reference.md (worked triage examples for real regression vs test bug vs flaky judge) or remove the 'Additional resources' section — the link currently points to a file that does not exist in the bundle.

Delete the duplicated vitest command under "To reproduce judge graders locally" in Rule 0, or replace it with the genuinely different command it was presumably meant to hold.

Align step numbering/references: Step 4 says 'Re-run the single scenario (Step 0 command)' but the section is titled 'Rule 0', and the scenario-id placeholder is never tied to where scenario ids are listed (e.g., the scenarios/ directory or the -t match string).

DimensionReasoningScore

Conciseness

The body is dense and lean — file-path tables, a symptom→verdict→fix-target decision table, and bare commands with no explanation of concepts Claude already knows. Not a 5 because of one clear redundancy: the "To reproduce judge graders locally" block in Rule 0 repeats the exact same vitest command already given two lines above, and "Step 0 command" in Step 4 refers to a section labeled "Rule 0" (a small naming mismatch). Not a 3: there is no unnecessary explanation or padding anywhere else.

4 / 5

Actionability

Fully executable guidance throughout: a copy-paste-ready re-run command (`pnpm --filter @novu/agent-evals exec vitest run --config vitest.evals.config.ts -t <scenario-id>`), exact file paths for every grader layer (`src/suites/agent-onboarding/scenarios/<id>/graders.ts`, `catalog.ts`, `src/core/graders.ts`, `src/core/judge.ts`), the concrete pass threshold (`averages ≥ 0.8 (JUDGE_THRESHOLD)`, `UNKNOWN`→`skip` scores 1), a grader-kind decision table, and a fill-in-the-blanks output template. Not a 4: there are no gaps — the common triage case is covered end-to-end.

5 / 5

Workflow Clarity

The sequence is explicit and gated: Rule 0 (re-run 3–5× to rule out flakiness before changing anything) → identify grader kind → read RunResult evidence → classify via top-down decision table → apply one bounded fix and verify. Validation checkpoints are explicit at every stage: "Confirm the fix holds across the 3–5 re-runs and that no other scenario regressed" and running the synthetic unit tests when a grader is edited. This matches the top anchor (feedback loops, explicit validation, ordered first-match classification).

5 / 5

Progressive Disclosure

The body itself is well structured with clear sections, but the single external reference is broken: "For worked triage examples ... see [reference.md](reference.md)" points to a file that does not exist in the bundle (no references/ directory), so navigation dead-ends and the promised worked examples are unavailable. Not a 4: a clearly signaled but missing reference is a real navigation defect, not a minor organization gap; not a 2: the body is appropriately sized and organized rather than a monolith that should be split.

3 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An excellent description: it states the skill's concrete triage actions, enumerates the possible fix targets, and provides an explicit 'Use when...' clause with multiple natural trigger phrases, all in third person and without padding. The only minor gap is a few missing user-side synonyms for the trigger terms.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "Triage failing @novu/agent-evals scenarios to decide whether a failure is real or flaky, and whether to fix the playbook/prompt or the test (grader, tape, scenario, or judge)" — with comprehensive coverage of the decision space (real vs flaky, playbook vs test, and the four test-layer targets). Not the level below (4): there are no meaningful coverage gaps; the enumerated fix targets (grader, tape, scenario, judge) make the actions fully concrete.

5 / 5

Completeness

Both halves are explicit: the 'what' (triage a failing scenario and classify the failure as real vs flaky, and identify the fix target) and the 'when' via a concrete "Use when..." clause with three distinct trigger conditions. This matches the top anchor exactly; the 'when' is explicit with concrete trigger phrases, not merely implied (level 4).

5 / 5

Trigger Term Quality

Good natural trigger phrases: "when an agent-evals scenario fails", "when the user asks why an eval is red", "when deciding whether to fix the test or the prompt" — phrasings a user would plausibly say. Not a 5: a few common synonyms are missing (e.g., "flaky test", "eval failing in CI", "failing scenario in the eval suite"). Not a 3: the coverage clearly exceeds 'some relevant keywords'.

4 / 5

Distinctiveness Conflict Risk

The niche is narrow and product-specific (@novu/agent-evals scenario triage with named layers like grader, tape, judge), so it is clearly distinguishable from generic testing/debugging skills and carries minimal conflict risk. Level 4 would require some overlap with closely related skills, which the strong domain binding avoids.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 missing

Warning

Total

15

/

16

Passed

Repository
novuhq/novu
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.