CtrlK
BlogDocsLog inGet started
Tessl Logo

pdf-gen-fallback

Reliable PDF generation using shell-based Python execution when sandbox fails

60

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./benchmarks/gdpval/skills/pdf-gen-fallback/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured and actionable, providing executable commands across a clear five-step workflow with verification, though it carries minor redundancy and an implicit rather than explicit error-recovery feedback loop.

Suggestions

Add an explicit validate->fix->retry feedback loop (e.g., 'If Step 5 shows the PDF is missing or invalid, re-run with stderr captured and adjust the command') to strengthen workflow clarity.

Remove or condense the 'Complete Example Workflow' section since it duplicates the Step-by-Step procedure, improving token efficiency.

Make Step 3's error detection actionable by giving a concrete command or check (e.g., capturing and grepping stderr) rather than only a list of indicators to watch for.

DimensionReasoningScore

Conciseness

The body is mostly efficient with short section intros and no concept-explaining padding, but the 'Complete Example Workflow' section restates the preceding steps, introducing minor redundancy that keeps it below a 5.

4 / 5

Actionability

Provides concrete, mostly copy-paste-ready run_shell and python -c commands for library checks, generation, and verification, with minor gaps such as the descriptive (non-command) error-detection list in Step 3 and the illustrative 'Complete Example Workflow'.

4 / 5

Workflow Clarity

A clear five-step sequence with a verification checkpoint (Step 5) and an error-detection step (Step 3), but the validate->fix->retry recovery loop is implicit rather than an explicit feedback loop, so it does not reach a 5.

4 / 5

Progressive Disclosure

Well-organized into clear sections (When to Use, Step-by-Step, Error Handling, Library Alternatives, Common Pitfalls) with no nested or broken references and no bundle files present; the modest length and step/example redundancy are minor organization gaps versus a clean 5.

4 / 5

Total

16

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly conveys a specific niche — sandbox-failure PDF generation fallback — and includes an explicit trigger condition, but its trigger vocabulary leans on internal jargon and omits natural user phrasing and synonyms.

Suggestions

Add natural user-facing trigger terms and synonyms (e.g., 'create PDFs', 'generate reports', '.pdf', 'documents') so the description matches how users actually phrase the request.

Expand the 'when' clause to cover the broader use case (e.g., 'Use when generating PDFs, reports, or structured documents, especially when execute_code_sandbox returns opaque errors') rather than only the sandbox-failure condition.

List a couple more concrete actions (e.g., 'generate PDFs and reports, verify output, fall back to shell execution') to lift specificity above the single-mechanism description.

DimensionReasoningScore

Specificity

Names the domain ('PDF generation') and one concrete action mechanism ('shell-based Python execution', 'sandbox fails'), but does not enumerate multiple generation actions, fitting the '1-2 concrete actions, not comprehensive' anchor.

3 / 5

Completeness

Answers both 'what' (reliable PDF generation via shell-based Python) and 'when' (when sandbox fails); the explicit 'when sandbox fails' trigger satisfies the 'Use when...' requirement, though the trigger is narrow rather than broadly applicable.

4 / 5

Trigger Term Quality

Contains relevant keywords like 'PDF generation' and 'sandbox fails', but relies on Claude-internal jargon ('shell-based Python execution', 'sandbox') and omits common natural variations like 'create PDFs', 'documents', 'reports', or '.pdf'.

3 / 5

Distinctiveness Conflict Risk

The fallback-trigger framing ('when sandbox fails') carves a fairly distinct niche with minimal conflict risk against general PDF skills, though it still shares the PDF-generation domain with closely related skills.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUDS/OpenSpace
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.