CtrlK
BlogDocsLog inGet started
Tessl Logo

paperclip-evals

Choose, inspect, validate, and report Paperclip Runner or Product E2E evaluations while preserving evidence, provenance, cost, and failure classification.

65

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/paperclip-evals/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a tight, actionable routing-and-safety guide with clear sequencing and well-signaled references; it could be slightly stronger with explicit error-recovery loops and a quick-start cell of commands.

DimensionReasoningScore

Conciseness

Lean and dense; it assumes Claude's competence (no explaining what evals/Daytona/protocols are) and every line carries routing or safety guidance, with no padding.

5 / 5

Actionability

Provides concrete, runnable commands (git rev-parse --show-toplevel, pnpm test:e2e:runner:typecheck/unit/--list) and specific file paths to read, with only minor gaps in fully copy-paste-ready sequences.

4 / 5

Workflow Clarity

Clear Discover -> Route -> Work safely sequence with explicit validation/authorization checkpoints (credential-free first, authorized + configured before a live cell); not capped at 3 because validation is present, though error-recovery feedback loops are only implicit.

4 / 5

Progressive Disclosure

Well-organized sections with clearly signaled, one-level-deep references to repo docs (doc/evals.md, README/FIXTURES/SECURITY/EVERYDAY-WORKFLOWS) and sibling skills; no bundle files ship with the skill, so references resolve to repo paths.

4 / 5

Total

17

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, distinct, and action-rich, but it lacks an explicit when-to-use trigger clause, which caps its completeness score.

Suggestions

Append a "Use when ..." clause naming the request types that should trigger this skill (e.g., selection, interpretation, evidence, history, or a live eval run) so the description answers both what and when.

Add a couple of natural synonyms users might say (e.g., "Paperclip evals", "E2E evals") to round out trigger term coverage.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ("Choose, inspect, validate, and report") plus concrete preserved attributes (evidence, provenance, cost, failure classification), giving comprehensive coverage of the eval workflow.

5 / 5

Completeness

Clearly states what the skill does but omits any "Use when..." trigger guidance in the description; per the rubric a missing when-clause caps completeness at 3.

3 / 5

Trigger Term Quality

Strong domain keywords ("Paperclip Runner", "Product E2E evaluations", "evidence", "provenance", "cost", "failure classification") a user would naturally say, though a few natural phrasings/synonyms are absent.

4 / 5

Distinctiveness Conflict Risk

Scoped to a clearly distinct niche (Paperclip Runner / Product E2E evaluations) with minimal risk of triggering for unrelated skills.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 2 suspicious

Warning

Total

15

/

16

Passed

Repository
paperclipai/paperclip
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.