CtrlK
BlogDocsLog inGet started
Tessl Logo

dogfood

Systematically explore and test a mobile app on iOS/Android with agent-device to find bugs, UX issues, and other problems. Use when asked to dogfood, QA, exploratory test, find issues, bug hunt, or test this app on mobile.

70

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A disciplined router skill: it stays tiny, gates setup with an explicit version check and failure path, and deliberately delegates volatile details to `agent-device help dogfood`. The gaps are the pseudocode-style loop line, an inline time-sensitive version floor, and underspecified evidence-capture steps in the body itself.

Suggestions

Expand the arrow-notation loop ("open app -> snapshot -i + screenshot -> ...") into a numbered step list with one concrete example command per step (e.g., the actual open/snapshot invocations), keeping `help dogfood` as the source of full detail.

Move the "agent-device >= 0.14.0" floor into a short, clearly-labeled compatibility/version-note section (or fold it into the `help dogfood` guidance step) so time-sensitive details are quarantined from the evergreen workflow.

State the minimal in-body evidence expectations for an issue (e.g., what a finding record must contain) so report shape does not depend entirely on the deferred CLI help.

DimensionReasoningScore

Conciseness

The body is lean with no concept explanations Claude already knows — every line is setup gating, commands, or workflow. However, the time-sensitive version floor "Require `agent-device >= 0.14.0`" sits inline rather than in an old-patterns/deprecated section, which the guidelines say should penalize conciseness; otherwise this would be a 5.

4 / 5

Actionability

Quotes: "agent-device --version" and "agent-device help dogfood" — two executable commands with explicit failure handling. But the core workflow is arrow notation ("open app -> snapshot -i + screenshot -> explore flows -> capture evidence per issue -> close") rather than copy-paste commands, so it is 'mostly executable with minor gaps' (level 4), not fully copy-paste ready (level 5). The deferral of exact commands to `help dogfood` is explicitly justified, which keeps this above level 3.

4 / 5

Workflow Clarity

Quotes: "If that fails, stop and tell the user to expose a trusted `agent-device` binary" and "If older, stop and tell the user to upgrade" — a clear setup→guidance→loop sequence with an explicit validation checkpoint and failure path. It stays at level 4 rather than 5 because the in-loop steps ("explore flows", evidence capture, report shape) are delegated without inline checkpoints; the destructive/batch validation cap does not apply since this is read-only exploration.

4 / 5

Progressive Disclosure

The body is ~27 lines with no bundle files, well-organized into router intro, version gate, guidance, and loop sections; per the simple-skill exception, an under-50-line skill with no need for external references can score 5 with well-organized sections. Its single external reference ("agent-device help dogfood") is one level deep and clearly signaled.

5 / 5

Total

17

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a strong example: it states what the skill does (exploratory mobile app testing with agent-device to find bugs and UX issues), scopes the platforms, and gives an explicit, synonym-rich 'Use when' trigger list. Its only weakness is the vague "other problems" phrase and minor coverage gaps in the action list.

DimensionReasoningScore

Specificity

Quotes: "Systematically explore and test a mobile app on iOS/Android with agent-device to find bugs, UX issues" — names the domain (iOS/Android mobile testing via agent-device) and several concrete actions (explore, test, find bugs, find UX issues). "Other problems" is generic padding and coverage omits evidence capture/reporting, so it fits the 'several specific actions; minor gaps' anchor rather than the comprehensive level 5.

4 / 5

Completeness

Quotes: "explore and test a mobile app on iOS/Android with agent-device to find bugs, UX issues" (clear what) and "Use when asked to dogfood, QA, exploratory test..." (explicit when with concrete trigger phrases). Both questions are answered explicitly, matching the level-5 anchor exactly.

5 / 5

Trigger Term Quality

Quotes: "Use when asked to dogfood, QA, exploratory test, find issues, bug hunt, or test this app on mobile" — comprehensive coverage of natural phrasings with multiple synonyms (dogfood, bug hunt, exploratory test, test this app). Matches the level-5 'comprehensive coverage of natural terms including synonyms' anchor; file extensions are not applicable to this domain, and no common variation is missing.

5 / 5

Distinctiveness Conflict Risk

Quotes: "test a mobile app on iOS/Android with agent-device" with triggers like "dogfood" and "bug hunt" — a clear niche (runtime mobile QA via a specific CLI) with distinct trigger terms and minimal overlap with other skills.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
callstack/agent-device
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.