CtrlK
BlogDocsLog inGet started
Tessl Logo

daytona-flow-validator

do E2E tests, validate feature, prove it works, pass/fail, test evidence, screenshots, CDP assertions. Daytona validation loop for real app behavior with repair before declaring success.

62

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./.opencode/skills/daytona-flow-validator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a strong, highly actionable validation playbook with explicit sequenced workflows, validation checkpoints, and repair feedback loops, all tailored to domain specifics Claude would not already know. Its only real weakness is that everything lives in one ~250-line file with no bundled reference files to split out the longer automation recipes.

Suggestions

Move the longer bash recipes (Linux Desktop Automation, native picker patterns) into a references/ file and link to it from SKILL.md to deepen progressive disclosure.

Trim a few rationale sentences (e.g. 'A recording that cannot be understood without terminal logs...') to tighten conciseness further.

Resolve the inline cross-references: clarify whether 'prove-a-pr', 'write-a-spec', and 'daytona-recording-artifacts' are sibling skills or files, and link them explicitly.

DimensionReasoningScore

Conciseness

The body is largely lean operational guidance specific to the Daytona/OpenWork domain (sandbox commands, CDP patterns, file paths) that Claude would not already know, with only minor prose that could be trimmed (e.g. rationale sentences in the demo standard).

4 / 5

Actionability

It provides copy-paste-ready bash commands, a JS paste-composer snippet, concrete tool calls (browser_snapshot/click/fill, daytona exec), and a fill-in validation-loop template that cover the common cases.

5 / 5

Workflow Clarity

Multi-step processes are explicitly sequenced with validation checkpoints: the Core Rule's 5-step observe/act/assert loop, the 5-step Repair Loop with feedback/retry, Failure Handling, and a pass/fail/incomplete Final Verdict definition.

5 / 5

Progressive Disclosure

The skill is a single well-sectioned file with clear headers and named cross-references (prove-a-pr, write-a-spec, daytona-recording-artifacts), but no bundle files exist and some inline material (Linux desktop automation, Lexical composer) could arguably live in one-level-deep reference files.

4 / 5

Total

18

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description conveys a clear, specific set of validation actions and a distinct Daytona/CDP niche, but it lacks an explicit 'Use when' trigger clause and relies on technical jargon over natural user phrasing. These gaps hold it at a mid-range score despite good specificity and distinctiveness.

Suggestions

Add an explicit trigger clause, e.g. 'Use when verifying a Daytona Electron or browser flow actually works before reporting success.'

Soften jargon with natural phrasings users would say ('test the app end-to-end', 'prove the flow works') alongside 'CDP assertions'.

Rewrite the comma-fragment list into clean third-person verb phrases ('Runs E2E tests, validates features, captures evidence') to improve readability and specificity.

DimensionReasoningScore

Specificity

The description lists several concrete actions ('do E2E tests, validate feature, prove it works, pass/fail, test evidence, screenshots, CDP assertions, ... repair before declaring success'), giving broad coverage of the domain, though the fragmentary comma-list leaves minor gaps in framing.

4 / 5

Completeness

The 'what' is clearly stated, but there is no explicit 'Use when...' clause or equivalent trigger guidance, so per the rubric completeness is capped at 3.

3 / 5

Trigger Term Quality

It contains relevant terms ('E2E tests', 'validate feature', 'pass/fail', 'screenshots', 'CDP assertions') but leans on technical jargon and misses more natural user phrasings and synonyms a person would spontaneously say.

3 / 5

Distinctiveness Conflict Risk

The 'Daytona validation loop for real app behavior with repair' framing carves a fairly distinct niche (CDP/Electron/browser flow validation) with only minor overlap risk against adjacent testing skills.

4 / 5

Total

14

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
different-ai/openwork
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.