CtrlK
BlogDocsLog inGet started
Tessl Logo

graceful-tool-failure

Report tool failures transparently and attempt minimal viable output with disclaimers

51

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./benchmarks/gdpval/skills/graceful-tool-failure/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a well-structured, actionable workflow with concrete templates and code, though it carries some redundancy between its template and principles sections. It is a solid, usable skill body with minor tightening opportunities.

Suggestions

Collapse the duplicate response templates (the Step 2/3 inline blocks and the "Response Template for Failed Tasks" example) into a single canonical template to improve conciseness.

Define or stub the Python helper functions (run_shell, extract_with_pymupdf, extract_with_pdfplumber) so the fallback example is copy-paste executable, lifting actionability toward 5.

Merge "Key Principles" and "Anti-Patterns to Avoid" into one consolidated do/don't list to remove overlapping content.

DimensionReasoningScore

Conciseness

The body is mostly efficient and free of concept-over-explanation, but the two overlapping response templates and the redundant "Key Principles" plus "Anti-Patterns" sections could be tightened, matching the score-3 anchor.

3 / 5

Actionability

Concrete markdown templates, a correct-vs-incorrect phrasing example, and a Python fallback snippet give mostly executable guidance; the only minor gap is the Python example's undefined helper functions (run_shell, extract_with_pymupdf), fitting score 4 over 5.

4 / 5

Workflow Clarity

A clear four-step sequence (report, attempt minimal output, avoid false success, next steps) with a "When to Apply" entry gate is well-structured; no internal validation loop is needed because this is a reporting endpoint rather than a destructive/batch operation, so score 4 rather than 5.

4 / 5

Progressive Disclosure

Content is well-organized into clearly headed sections with no external references needed; because the body exceeds 50 lines it does not qualify for the under-50-line score-5 exception, so it sits at the good-structure score-4 anchor.

4 / 5

Total

15

/

20

Passed

Description

42%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states a clear purpose and two concrete actions but lacks any explicit trigger guidance, which limits completeness and trigger-term quality. It is a reasonable but incomplete skill description that would benefit from a "Use when..." clause.

Suggestions

Add an explicit trigger clause such as "Use when a core tool has failed repeatedly (3+ attempts) and blocks task completion" to answer the "when" and lift completeness.

Include natural user-facing trigger terms (e.g., "tool keeps failing", "can't extract", "tool error") alongside the technical phrasing to improve trigger-term quality.

Sharpen distinctiveness by naming the specific scenario (repeated tool failures blocking a deliverable) rather than the generic "report failures" framing.

DimensionReasoningScore

Specificity

The description names the domain (tool failures) and two concrete actions — "Report tool failures transparently" and "attempt minimal viable output with disclaimers" — but coverage is not comprehensive, matching the score-3 anchor.

3 / 5

Completeness

A clear "what" is present (report failures, produce disclaimed minimal output) but there is no "Use when..." clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

Only "tool failures" reads as a natural user phrase; "minimal viable output" and "disclaimers" are process jargon, and common variations a user would actually say are missing, fitting the score-2 anchor.

2 / 5

Distinctiveness Conflict Risk

The failure-reporting niche is somewhat specific, but as a broad meta/process skill it could still overlap with general error-handling or reporting skills, matching the score-3 anchor rather than the clearly-distinct score-4 anchor.

3 / 5

Total

11

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUDS/OpenSpace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.