CtrlK
BlogDocsLog inGet started
Tessl Logo

run-ebitengine-app-headless

Use this skill to run an Ebitengine app (anything implementing ebiten.Game) headlessly and programmatically — no visible window and no changes to the app's source — to test, debug, screenshot, or otherwise exercise it. It launches the app as a vmhost guest, drives it through its ticks faster than real time, injects keyboard / mouse / touch / gamepad / text input, reads the rendered frame back as pixels to assert on or dump as a PNG, and observes the audio the app plays. Use when you need input injection, deterministic multi-tick runs, golden-image checks, or audio assertions — for example to reproduce an input-dependent bug, verify behavior after a sequence of clicks or key presses, check that the right sound plays, capture a single rendered frame, or let an AI agent drive an app end-to-end.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable skill body: copy-paste-ready commands and code, clearly sequenced recipe steps, and unusually honest verification guidance covering failure diagnosis, flaky launches, and golden-image pitfalls. Its main weaknesses are the dense single-file layout — several advanced deep dives would be better as one-level-deep references — and minor redundancy plus an upfront pinned-commit note that adds time-sensitive information.

Suggestions

Split the self-contained deep dives ("Observing audio", the gamepad touchpad semantics under recipe step 2, and "In-repo alternative (Go test)") into one-level-deep reference files (e.g. references/audio.md, references/gamepads.md) signaled from the main body, slimming SKILL.md toward a lean overview.

Trim the overlap between "How the driver drives the guest" and the recipe's AdvanceTicks/WaitFrame material, and consider moving the pinned commit (7de1780bd, 2026-09-27) into a short compatibility/versions note so time-sensitive details don't age the main body.

In the golden-compare workflow, show the exact byte-comparison command (e.g. cmp -s baseline.png /tmp/frame.png with the existing-file pre-check) so the verification loop is copy-paste ready like the rest of the skill.

DimensionReasoningScore

Conciseness

The body is efficient and assumes Claude's competence — it explains only the non-obvious exp/vmhost mechanics (host/guest split, endpoint wiring, snapshot semantics, audio stream lifecycle) rather than background Go or Ebitengine concepts. It falls at anchor 4 rather than 5 because of a pinned time-sensitive commit ("Targets ebiten commit 7de1780bd (2026-09-27)" outside any old-patterns section) and some redundancy between the recipe and the 'How the driver drives the guest' section that could be trimmed.

4 / 5

Actionability

Guidance is fully executable: copy-paste bash commands with all flags ("go run ./skills/run-ebitengine-app-headless/_driver -pkg ./examples/rotate -ticks 60 -out /tmp/frame.png"), complete Go injector snippets, an exact pixel-offset formula ("the center pixel of that physical w*h image is at 4*((h/2)*w + w/2)"), and a concrete failure-capture command ("> /tmp/run.log 2>&1 || cat /tmp/run.log"). The few placeholders (pageCount, nextButtonX) are explicitly framed as template adaptation points.

5 / 5

Workflow Clarity

The recipe (steps 1–4) is clearly sequenced from no-input capture through scripted input, multi-state snapshots, and inspection, and the 'Verifying a run' section supplies explicit validation checkpoints: judge success by PNG presence rather than piped exit status, save full output to a log before diagnosing, confirm both files exist before byte-comparing, plus retry-once guidance for flaky launches and a safe git-worktree pattern for golden baselines. These are genuine feedback loops for a fragile operation.

5 / 5

Progressive Disclosure

Structure is good: well-labeled sections, an out-of-line driver asset ("A host driver lives at _driver/main.go") positioned as a template, and clear internal anchor links ([Ticks and TPS](#ticks-and-tps), [Verifying a run](#verifying-a-run)). It stops at anchor 4 rather than 5 because the ~365-line body inlines several self-contained deep dives (audio stream inspection, gamepad touchpad semantics, the in-repo Go-test alternative) that would fit one-level-deep reference files, leaving the main file dense.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: a precise scope statement, a comprehensive third-person list of concrete capabilities, and an explicit 'Use when…' clause with varied natural trigger phrases and concrete example scenarios. It is long but every clause is a distinct capability or trigger, so the length is dense rather than padded.

DimensionReasoningScore

Specificity

The description lists multiple concrete, specific actions — "launches the app as a vmhost guest", "drives it through its ticks faster than real time", "injects keyboard / mouse / touch / gamepad / text input", "reads the rendered frame back as pixels to assert on or dump as a PNG", "observes the audio the app plays" — with comprehensive coverage and no vague filler, matching the anchor 'Lists multiple specific concrete actions; comprehensive coverage' rather than the anchor below (which requires minor gaps in coverage).

5 / 5

Completeness

It explicitly answers both questions: a detailed third-person 'what' ("It launches the app as a vmhost guest… injects… reads the rendered frame back… observes the audio") followed by an explicit trigger clause — "Use when you need input injection, deterministic multi-tick runs, golden-image checks, or audio assertions — for example to…" — with concrete trigger phrases, exactly the anchor-5 pattern.

5 / 5

Trigger Term Quality

Natural user phrasing is comprehensively covered: "test, debug, screenshot", "input injection, deterministic multi-tick runs, golden-image checks, or audio assertions", plus scenario-level triggers like "reproduce an input-dependent bug", "verify behavior after a sequence of clicks or key presses", "check that the right sound plays", and "capture a single rendered frame". This matches the comprehensive-with-synonyms anchor; no natural term is obviously missing.

5 / 5

Distinctiveness Conflict Risk

It opens with a sharply bounded niche — "run an Ebitengine app (anything implementing ebiten.Game) headlessly and programmatically — no visible window and no changes to the app's source" — and every trigger is Ebitengine- or headless-testing-specific, so it is clearly distinguishable from generic run/test/screenshot skills with minimal conflict risk.

5 / 5

Total

20

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

relative_links

Relative link issues: 5 missing, 5 deeper-than-1-level

Warning

Total

14

/

16

Passed

Repository
hajimehoshi/ebiten
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.