CtrlK
BlogDocsLog inGet started
Tessl Logo

terminal-bench-loop

Run one Terminal-Bench task through a bounded Paperclip smoke/diagnosis/fix loop. Use when asked to drive Terminal-Bench until it passes, rerun a Terminal-Bench loop, or iterate with board-gated fixes and diagnosis.

68

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An unusually rigorous operational contract: the procedure, stop rules, and checklist give excellent workflow clarity and the guidance is specific down to exact field names and idempotency keys. Its weaknesses are repetition of the same liveness/topology and worktree rules across sections, and a monolithic single-file structure whose one canonical external reference (doc/execution-semantics.md) is not part of the bundle.

Suggestions

Consolidate the typed-waiter/in_review topology rule and the worktree/PAPERCLIPAI_CMD rule into single authoritative sections and reference them elsewhere in one line, instead of restating them in Issue topology, Steps 3/6/7/8, Worktree rule, Liveness rule, and Pitfalls.

Split the detailed dispatch-runner-config input spec, the Pitfalls list, and the Deterministic smoke details into reference files under references/ so SKILL.md stays a lean overview with one-level-deep, clearly signaled pointers.

Resolve the doc/execution-semantics.md reference: either ship it inside the skill bundle (e.g. references/execution-semantics.md) or state inline where the document lives, since it is declared canonical reading for every loop start but is absent from the skill.

DimensionReasoningScore

Conciseness

The body never explains concepts Claude already knows, but it restates the same rules many times: the typed-waiter/`in_review` topology rule appears in "Issue topology", Steps 6, 8, the "Liveness rule", and "Pitfalls", and the worktree/`PAPERCLIPAI_CMD` rule appears in five sections. It is mostly efficient but could be tightened considerably by stating each rule once, matching the anchor-3 example better than anchor 4's "minor instances".

3 / 5

Actionability

Guidance is fully concrete and executable for its domain: exact statuses and field names (`blockedByIssueIds`, `inheritExecutionWorkspaceFromIssueId`), an exact idempotency key format (`confirmation:{iterationIssueId}:plan:{revisionId}`), a `continuationPolicy` value, an exact issue title template, and a runnable verification command (`pnpm smoke:terminal-bench-loop-skill`). Per the rubric's instruction-only note, the absence of code is not penalized when the guidance is this actionable.

5 / 5

Workflow Clarity

Steps 0-9 are clearly sequenced with five explicit terminal outcomes (Step 5), explicit stop rules (Step 9), a verification checklist, and a built-in feedback loop (run -> diagnose -> board confirm -> implement -> QA -> rerun -> re-diagnose). Every risky transition has a validation checkpoint, matching the anchor-5 example.

5 / 5

Progressive Disclosure

The body is a ~230-line monolith with good section headers, but material that clearly belongs in separate files is inlined (detailed dispatch-config input spec, the full pitfalls list, the smoke script details), and the canonical reference `doc/execution-semantics.md` points to a file outside the skill bundle (no references/ directory ships with it), which is not clearly resolved. This is "some structure but could be better organized" rather than the well-signaled one-level-deep references of anchor 5.

3 / 5

Total

16

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A tight, third-person description with an explicit what and a well-formed "Use when..." clause carrying several natural trigger phrasings. It is highly distinct within its niche; the only room for improvement is slightly fuller enumeration of the concrete actions and a few more natural synonyms.

DimensionReasoningScore

Specificity

"Run one Terminal-Bench task through a bounded Paperclip smoke/diagnosis/fix loop" names the domain plus several concrete actions (smoke, diagnosis, fix) with only minor coverage gaps (board gating and worktree continuity live only in the body). It lists more than the 1-2 actions of the anchor-3 example but is not the comprehensive multi-action enumeration of the anchor-5 example.

4 / 5

Completeness

It explicitly answers both: the "what" ("Run one Terminal-Bench task through a bounded Paperclip smoke/diagnosis/fix loop") and the "when" ("Use when asked to drive Terminal-Bench until it passes, rerun a Terminal-Bench loop, or iterate with board-gated fixes and diagnosis") with concrete trigger phrases, matching the anchor-5 example.

5 / 5

Trigger Term Quality

"drive Terminal-Bench until it passes", "rerun a Terminal-Bench loop", and "iterate with board-gated fixes and diagnosis" are natural phrasings users would say, covering the main verb variations. A few natural synonyms (e.g. "Terminal-Bench task", "bench task name") are missing, and "board-gated" is more jargon than user language, so it sits at good-but-not-comprehensive coverage.

4 / 5

Distinctiveness Conflict Risk

"Terminal-Bench" and "Paperclip" are niche-specific proper nouns with distinct, unambiguous triggers; the description carves a clear loop-driving niche that no generic skill would accidentally match.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
paperclipai/paperclip
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.