CtrlK
BlogDocsLog inGet started
Tessl Logo

run-evals

do e2e tests, run e2e, validate feature, prove it works, PR proof, frame proof, pnpm evals. Launches iPolloWork on Daytona or local Electron and runs the coded eval flows via CDP. Launch + run mechanics; the proof loop itself is the fraimz skill.

68

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a tight, fully executable runbook with clear sequencing and an explicit CDP verification checkpoint plus error-recovery guidance. Structure is good; the only minor gap is reliance on external skills/docs rather than bundled reference files for deeper detail.

DimensionReasoningScore

Conciseness

Lean and efficient throughout: command snippets and brief rationale only, assuming Claude's competence with Daytona, Electron, and CDP, with no conceptual padding.

5 / 5

Actionability

Fully executable, copy-paste-ready commands cover the common cases (Daytona path, verification curl, running flows, local fallback, teardown) with concrete flags and URLs.

5 / 5

Workflow Clarity

Clear sequence with an explicit validation checkpoint ('curl -fsS <CDP_URL>/json/list') and a feedback loop (inspect /tmp/electron.log for the real success marker) for error recovery.

5 / 5

Progressive Disclosure

Well-organized into clearly signaled sections with one-level-deep references to the fraimz skill and evals/README.md, but detail that could live in bundled reference files is instead deferred to external skills.

4 / 5

Total

19

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific and full of natural trigger terms, and it carves out a clear niche with an explicit boundary to a sibling skill. Its main weakness is the absence of an explicit 'Use when...' trigger clause, which caps its completeness.

Suggestions

Add an explicit 'Use when...' clause, e.g. 'Use when you need to run coded e2e evals against a real iPolloWork build and produce PR proof.'

Trim the redundant comma-separated trigger list ('do e2e tests, run e2e, validate feature') to the most natural phrases to reduce buzzword density.

Clarify what 'frame proof' / 'fraimz' means or drop these internal jargon terms so the trigger guidance is intelligible without the sibling skill loaded.

DimensionReasoningScore

Specificity

Lists several concrete actions ('Launches iPolloWork on Daytona or local Electron and runs the coded eval flows via CDP') rather than vague language, with only minor coverage gaps.

4 / 5

Completeness

The 'what' is explicit but the 'when' is only weakly implied via trigger terms; there is no explicit 'Use when...' clause, which caps completeness at 3 per the rubric guideline.

3 / 5

Trigger Term Quality

Includes natural phrases a user would actually say ('do e2e tests', 'run e2e', 'validate feature', 'prove it works', 'PR proof') plus the command term 'pnpm evals', though a few synonyms are missing.

4 / 5

Distinctiveness Conflict Risk

Has a clear niche (launch + run mechanics) and explicit boundary with the fraimz skill, leaving only minor overlap risk with that closely related sibling skill.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Devin-AXIS/iPolloWork
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.