CtrlK
BlogDocsLog inGet started
Tessl Logo

nemoclaw-maintainer-audit-e2e-assertions

Triage, diagnose, debug, or fix failing or flaky NemoClaw E2E tests. Trace every assertion and downstream gate before a repair push or rerun. Excludes status-only and dispatch-only requests.

66

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A disciplined, dense process skill with an explicit pre-push checklist, feedback loops, and a well-specified ledger artifact. The gaps are minor: no worked example of a completed ledger row or assertion ID, and some rules are stated more than once.

Suggestions

Include one filled-in example ledger row (with a sample assertion ID, expected condition, traced inputs, and evidence status) so the required artifact is unambiguous from the spec alone.

Deduplicate rules restated across sections (e.g., completion-boundary conditions in "Maintain the Assertion Ledger" vs. "Finish the Review Before the Next Push or Run") by stating each once and cross-referencing.

DimensionReasoningScore

Conciseness

The body is dense imperative instruction with no explanations of concepts Claude already knows, but rules are restated across sections (completion-boundary conditions in "Maintain the Assertion Ledger" reappear in "Finish the Review"; prohibition on weakening assertions appears twice). This is anchor 4 (efficient, minor instances that could be trimmed) rather than 5 (every token earns its place).

4 / 5

Actionability

Concrete directive guidance throughout: a ledger table with six specified columns, five enumerated evidence-status definitions, and specific mechanics like "Read large logs in sequential chunks" and stable-ID assignment rules. For an instruction-only skill this is mostly executable, but no filled example ledger row or example assertion-ID format is given, keeping it below anchor 5.

4 / 5

Workflow Clarity

A clear sequence (establish failure, enumerate contract, trace predicates, maintain ledger, batch repairs, finish) with an explicit pre-push checklist ("establish all of these conditions") and feedback loops ("A failure requires new causal evidence and an update to all affected rows before another repair push"; re-inspection of affected callers after edits). This matches anchor 5: explicit validation steps, error-recovery loops, and a checklist for a complex process.

5 / 5

Progressive Disclosure

Sections are well organized with descriptive headers and one clearly signaled external reference (the shared writing/review contract); no bundle files (references/, scripts/, assets/) exist to split content into, and the single-file process holds together. The ledger column format and evidence-status taxonomy could arguably live in a reference file, so it fits anchor 4 (good structure, minor organization gaps) rather than 5.

4 / 5

Total

17

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A tight, well-targeted description for a named product domain with explicit boundary exclusions. Its main weaknesses are near-synonymous verb lists instead of distinct capabilities and missing common trigger synonyms such as CI failure or timeout.

Suggestions

Replace the near-synonymous verb run ("Triage, diagnose, debug, or fix") with distinct capabilities, e.g. "Audits every assertion in failing or flaky NemoClaw E2E tests and builds a per-assertion evidence ledger before any repair push or rerun."

Add natural trigger variations such as "CI failure", "test timeout", or "assertion failure" so the skill surfaces for users who do not say "E2E".

DimensionReasoningScore

Specificity

"Triage, diagnose, debug, or fix" plus "Trace every assertion and downstream gate" name several concrete actions in a narrow domain, but the four verbs are near-synonyms covering one capability cluster and the ledger deliverable is unmentioned. This sits at anchor 4 (several specific actions, minor coverage gaps) rather than 5 (comprehensive distinct actions) and clearly above 3 (only 1-2 actions).

4 / 5

Completeness

The "what" is clear (triage/diagnose/fix failing E2E tests, trace assertions and gates) and the "when" is addressed via "before a repair push or rerun" plus the explicit exclusion clause "Excludes status-only and dispatch-only requests", which constitutes equivalent trigger guidance so the 3-cap does not apply. It is not a 5 because there is no direct "Use when..." phrasing with concrete trigger phrases.

4 / 5

Trigger Term Quality

"failing or flaky", "E2E tests", "repair push", and "rerun" are phrases a user would naturally say when needing this skill. Common variations like "CI failure", "test timeout", or "assertion failure" are absent, so it fits anchor 4 (good coverage, a few natural terms missing) rather than 5.

4 / 5

Distinctiveness Conflict Risk

"NemoClaw E2E tests" is a clearly named niche, and the exclusion of status-only and dispatch-only requests explicitly separates it from a sibling execution/dispatch skill. This matches anchor 5 (clear niche, distinct triggers, minimal conflict risk).

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 suspicious

Warning

Total

15

/

16

Passed

Repository
NVIDIA/NemoClaw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.