CtrlK
BlogDocsLog inGet started
Tessl Logo

run-evals

do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals. Launches iPolloWork on Daytona or local Electron and runs the coded eval flows via CDP. Launch + run mechanics; the proof loop itself is the fraimz skill.

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplar of lean, command-driven instruction with a clearly sequenced workflow and genuine validation/feedback checkpoints (CDP endpoint check plus electron.log failure diagnosis). The only notable gaps are the undefined $SANDBOX variable and the absence of any reference files to offload secondary detail.

DimensionReasoningScore

Conciseness

Every section is dense operational instruction — commands with flags, one-line failure diagnostics, and brief boundary notes — with no explanation of concepts Claude already knows. This matches 'lean and efficient; assumes Claude's competence; every token earns its place'.

5 / 5

Actionability

Concrete, mostly copy-paste-ready commands throughout (daytona setup, test-on-daytona.sh with flags, curl CDP verification, pnpm evals invocations, teardown), matching 'mostly executable guidance with minor gaps'. Not a 5 because "$SANDBOX" in the recording and teardown commands is never defined or shown how to obtain, and the '.newtoken' argument to setup-daytona-secrets-volume.sh is unexplained.

4 / 5

Workflow Clarity

The sequence (prerequisites → Daytona launch → verify endpoint → run flows → recording → local fallback → teardown) has an explicit validation checkpoint ("curl -fsS \"<CDP_URL>/json/list\" # must include an iPolloWork page target") and an error-recovery feedback loop ("If it fails, inspect /tmp/electron.log — the real success marker is Chromium's DevTools listening on ws://..."), matching the 'clear sequence with explicit validation steps; feedback loops for error recovery' anchor.

5 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent) and the body is ~80 well-organized lines, so the simple-skill exception does not fully apply; external pointers (fraimz skill, evals/README.md, daytona-recording-artifacts) are clearly signaled. This matches 'good structure; most content appropriately placed; references mostly clear; minor organization gaps' — a 5 would require a leaner overview or one-level-deep reference files for the recording/troubleshooting detail.

4 / 5

Total

18

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concrete and highly distinctive, with a strong natural-language trigger list and an explicit boundary against the sibling fraimz skill. Its main weakness is that the trigger phrases are presented as a bare keyword list rather than an explicit 'Use when...' clause tying them to the capability.

Suggestions

Convert the leading keyword list into an explicit trigger clause, e.g. 'Use when asked to run e2e tests, validate a feature, prove a PR works, or run pnpm evals.'

Add one or two common trigger variations such as 'end-to-end' or 'smoke test' to broaden natural keyword coverage.

Clarify routing for overlapping triggers like 'prove it works' and 'PR proof' (e.g. 'use fraimz when the ask is the verdict/evidence loop; use this skill for launch and run mechanics').

DimensionReasoningScore

Specificity

"Launches iPolloWork on Daytona or local Electron and runs the coded eval flows via CDP" names concrete actions with platform and protocol specifics, matching the 'several specific actions; minor gaps' anchor. Not a 3 because the actions are far more concrete than 'names domain and 1-2 actions', and not a 5 because coverage is limited to launch/run while recording, teardown, and verification actions from the body are absent.

4 / 5

Completeness

The 'what' is explicit (launches iPolloWork on Daytona/local Electron, runs coded eval flows via CDP) and the trigger-phrase list provides explicit trigger guidance, matching 'both what and when; when could be more explicit'. Not a 5 because there is no 'Use when...' phrasing connecting the triggers to the capability — they read as a keyword tag list rather than an explicit usage condition.

4 / 5

Trigger Term Quality

The leading list "do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals" covers many natural phrases a user would say, matching 'good keyword coverage; a few natural terms missing'. Not a 5 because common variations like 'end-to-end', 'smoke test', or 'integration test' are missing; not a 3 because synonym coverage (PR proof, frame proof, prove it works) is well beyond 'some relevant keywords'.

4 / 5

Distinctiveness Conflict Risk

Highly niche vocabulary (iPolloWork, Daytona, CDP, pnpm evals) gives a clear niche, matching 'mostly distinct; minor overlap risk'. Not a 5 because triggers like "prove it works" and "PR proof" overlap with the sibling fraimz skill, creating some routing ambiguity even though the boundary note "the proof loop itself is the fraimz skill" mitigates it.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Devin-AXIS/iPolloWork
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.