CtrlK
BlogDocsLog inGet started
Tessl Logo

boltz-check-status

Boltz job status and result recovery. Use when listing jobs, checking progress, resuming downloads, recovering results, or downloading an existing job ID. Not for starting new jobs.

71

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

80%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, token-efficient body with excellent progressive disclosure and real operational detail (pagination caps, output formats, per-host behavior). The main defect is an inconsistent mode-numbering scheme between the Workflow and Command Pattern sections, plus some duplicated guidance that could be consolidated.

Suggestions

Reconcile the mode numbering: the Workflow section numbers local-progress/list/retrieve/resume as Modes 1-4, while the Command Pattern comments and 'Always Do This' label list/retrieve/download as Modes 1-3 — adopt one scheme so 'run Mode 2 before Mode 3' is unambiguous.

Deduplicate repeated guidance: the download-status-first preference and the never-rerun-start rule each appear two to three times across Workflow, Always Do This, and Outputs — state each once in Always Do This and reference it.

Consider moving the host-runtime specifics (Codex heartbeat automation, write_stdin polling, foreground/yield rules) into a short reference file so SKILL.md stays a lean overview for all runtimes.

DimensionReasoningScore

Conciseness

Overall efficient — no explanations of concepts Claude already knows, and dense with non-obvious operational detail (streamed JSONL output, auto-pagination, idempotency_key capture). Not a 5 because guidance is repeated across sections ("prefer `download-status` before remote calls" appears in Workflow, Always Do This, and Outputs; "never run `start` again" appears at lines 19 and 73) and several host-specific sentences (Codex heartbeat/yield handling) run long enough to be distracting.

4 / 5

Actionability

Fully executable copy-paste-ready commands for every operation: all six resource `list` calls with `--limit`/`--format`/`head` caps, all six `retrieve` calls, and `download-status`/`download-results` invocations with concrete flags and a note to substitute placeholders with absolute paths. Not a 4 because the common cases (list, retrieve, resume) are each covered with complete, runnable commands plus expected output semantics.

5 / 5

Workflow Clarity

The sequence and decision rules are present (prefer `download-status` locally, retrieve before download to capture `idempotency_key`, probe resources until one succeeds, never re-run `start`), but the mode numbering is inconsistent: the Workflow section defines four modes (1=local progress, 2=list, 3=retrieve, 4=resume) while the Command Pattern and Always Do This sections use a different numbering (Mode 1=list, Mode 2=retrieve, Mode 3=download), leaving "Mode 1" and "Mode 2" ambiguous. Not a 4 because this labeling conflict is a real navigation gap, not a minor one; not a 2 because the underlying steps and checkpoints are otherwise clearly sequenced and the skill is read-only, so the destructive/batch validation cap does not apply.

3 / 5

Progressive Disclosure

The body is a well-organized overview with clearly signaled, one-level-deep references: "Read [references/resume.md] before recovering a dropped session..." and "Read [references/api.md] for per-resource `list` columns, `retrieve` fields..." — both files exist, contain the promised material, and reference only each other (no nested chains). Not a 4 because navigation is explicit about when each file is needed and the split (overview vs. per-resource detail vs. resume semantics) matches the anchor example exactly.

5 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A tight, third-person description that pairs a precise capability statement with an explicit "Use when" trigger list and a negative boundary clause. It is concise, uses natural trigger language, and is well-differentiated from a companion job-starting skill.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — "listing jobs, checking progress, resuming downloads, recovering results, or downloading an existing job ID" — which comprehensively covers the skill's stated scope of "job status and result recovery". Not a 4 because there is no noticeable coverage gap within that scope; each operation mode is named as a distinct action.

5 / 5

Completeness

Explicitly answers both questions: what ("Boltz job status and result recovery") and when ("Use when listing jobs, checking progress, resuming downloads, recovering results, or downloading an existing job ID"), plus an explicit negative boundary ("Not for starting new jobs"). Matches the score-5 anchor with concrete trigger phrases; not a 4 because the when-clause is already fully explicit, not merely adequate.

5 / 5

Trigger Term Quality

Good natural-keyword coverage: "listing jobs", "checking progress", "resuming downloads", "recovering results", "downloading", "job status", "job ID" — phrases a user would plausibly say. Not a 5 because it omits some natural variations/synonyms users might use (e.g., "check on my run", "fetch/get results", "finished yet"); not a 3 because the core trigger vocabulary is well covered beyond just domain naming.

4 / 5

Distinctiveness Conflict Risk

Clear niche (Boltz jobs) with distinct triggers, and the "Not for starting new jobs" exclusion explicitly separates it from a job-submission skill. Minimal conflict risk; not a 4 because the boundary is stated outright rather than only implied.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
openai/plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.