CtrlK
BlogDocsLog inGet started
Tessl Logo

run-evals

do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals. Launches iPolloWork on Daytona or local Electron and runs the coded eval flows via CDP. Launch + run mechanics; the proof loop itself is the fraimz skill.

74

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

100%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, highly actionable launch+run skill: concrete commands throughout, an explicit validation checkpoint with error-recovery branches, and well-signaled references deferring deeper detail to sibling docs and the fraimz skill.

DimensionReasoningScore

Conciseness

Lean and information-dense across ~85 lines — every section earns its place and it assumes Claude knows CDP/VNC/Electron rather than explaining them, matching the 'lean and efficient; every token earns its place' anchor rather than the padded level 2.

3 / 3

Actionability

Provides fully concrete, copy-paste-ready commands with real flags (e.g. 'bash .devcontainer/test-on-daytona.sh <branch-or-commit> --artifacts-volume', 'pnpm evals --flow <flow-id> --cdp-url <printed-electron-cdp-url>', 'daytona delete "$SANDBOX"'), not pseudocode.

3 / 3

Workflow Clarity

Clear sequence (Prerequisites → Preferred path → Verify → Run flows → Recording → Local fallback → Teardown) with an explicit validation checkpoint ('browser_list(...) must show an iPolloWork target') and error-recovery branches ('If it fails, inspect /tmp/electron.log', 'If the app shows the Welcome page...').

3 / 3

Progressive Disclosure

Organized into clear sections and signals one-level-deep detail rather than inlining it (e.g. 'see evals/daytona-flows.md Flow 1', 'see the fraimz skill and evals/README.md for the ctx.* API', 'Details: daytona-recording-artifacts'), keeping the body an overview; no bundle files exist to score against.

3 / 3

Total

12

/

12

Passed

Description

82%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A keyword-rich, concrete description that names the domain and multiple actions and de-conflicts from a sibling skill, but it lacks an explicit 'Use when...' trigger clause, which caps completeness at 2.

Suggestions

Add an explicit 'Use when...' clause (e.g., 'Use when the user asks to run e2e evals, validate a feature end-to-end, or produce PR/frame proof for iPolloWork') to lift completeness to 3.

Reframe the opening comma-separated keyword dump into a natural sentence so trigger terms read as guidance rather than a tag list.

Lead with the third-person capability sentence before the keyword list so the description scans as guidance first.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ('Launches iPolloWork on Daytona or local Electron and runs the coded eval flows via CDP', 'Launch + run mechanics') alongside a broad action vocabulary, matching the 'lists multiple specific concrete actions' anchor rather than the partial-coverage level 2.

3 / 3

Completeness

Clearly states what the skill does, but has no explicit 'Use when...' trigger clause — the when is only implied via the opening keyword dump, so per the guidelines completeness is capped at 2 and does not reach the explicit-trigger level 3.

2 / 3

Trigger Term Quality

Includes natural terms users would say — 'do e2e tests', 'run e2e', 'validate feature', 'prove it works', 'pnpm evals' — giving good coverage of common variations rather than just a few relevant keywords.

3 / 3

Distinctiveness Conflict Risk

Carves a clear niche ('Launch + run mechanics; the proof loop itself is the fraimz skill') and explicitly de-conflicts with the sibling skill, so it is unlikely to trigger for the wrong skill.

3 / 3

Total

11

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Devin-AXIS/iPolloWork
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.