CtrlK
BlogDocsLog inGet started
Tessl Logo

add-runner-eval

Add or extend a Paperclip Runner protocol evaluation definition, roster, assertion, or report fixture with provenance and narrow validation.

61

Quality

72%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/add-runner-eval/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is highly actionable with executable validation commands and concrete paths, and it structures references one level deep, but it relies on prose rather than headers/checklists and repeats the publication rules.

DimensionReasoningScore

Conciseness

The body is dense and assumes domain competence without explaining basics, but the reviewed-projection/publication rules are restated in two places, a minor redundancy that could be trimmed.

4 / 5

Actionability

It provides copy-paste-ready validation commands with real paths and flags, concrete env vars (PAPERCLIP_ROOT, PAPERCLIP_EVALS_ROOT), and specific authoring requirements, covering the common cases fully.

5 / 5

Workflow Clarity

A clear sequence is present with an explicit validation checkpoint ('Validate without provider calls first using the commands above'), but it is prose rather than a numbered checklist and lacks an explicit validate-fix-retry feedback loop.

4 / 5

Progressive Disclosure

No bundle files exist; references to external docs (doc/evals.md, runner-protocol-live-evals.md, sibling evals repo) are one level deep and clearly signaled, with content appropriately kept in one file, though section headers would improve navigation.

4 / 5

Total

17

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and clearly niche-scoped with concrete object types, but it omits any explicit 'Use when…' trigger guidance, which caps completeness and limits trigger-term coverage.

Suggestions

Add a 'Use when…' clause naming the natural trigger phrases (e.g. 'Use when adding or extending a Runner eval case, roster, assertion, or report fixture').

Include a couple of common synonyms or variations of the trigger terms so users can surface the skill with everyday phrasing.

Briefly state the boundary versus the sibling add-product-e2e-eval skill to reinforce distinctiveness in the description itself.

DimensionReasoningScore

Specificity

Names the Runner-eval domain and lists several concrete object types ('evaluation definition, roster, assertion, or report fixture') plus 'provenance and narrow validation', which is several specific actions with only minor coverage gaps.

4 / 5

Completeness

It gives a clear 'what' (add/extend eval definitions, rosters, assertions, fixtures) but includes no 'Use when…' trigger clause, so per the guideline completeness is capped at 3.

3 / 5

Trigger Term Quality

Terms like 'roster', 'assertion', 'report fixture', and 'provenance' are the natural vocabulary for this niche, but no synonyms or common variations are offered, matching the 'some relevant keywords but missing variations' anchor.

3 / 5

Distinctiveness Conflict Risk

The 'Paperclip Runner protocol evaluation' niche is tightly scoped with distinct triggers and minimal overlap risk with other skills.

5 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 1 suspicious

Warning

Total

15

/

16

Passed

Repository
paperclipai/paperclip
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.