CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-device-evidence

Records iOS/Android native MP4 evidence for test/repro flows extracted from an Expensify GitHub PR or issue. Use when the user asks to "record the flow for PR

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

High

Do not use without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality operational skill: fully executable commands, an explicitly ordered workflow with validation checkpoints at every phase, error-recovery feedback loops, and a clean one-level reference split (parsing rules, manifest schema, error matrix) with all referenced files present and non-nested. The only notable weakness is redundant restatement of platform/HybridApp/exit-code boundaries across multiple sections, which trims a little token efficiency without adding information.

Suggestions

Consolidate the boundary declarations: state the mWeb/Desktop `EXIT 4` and HybridApp `EXIT 7` gates once (e.g., in the exit-codes table) and reference them from Scope and 'Out of scope' rather than restating the full rationale in each place.

Remove the duplication between the 'Exit codes' table and references/error-handling.md — keep the table in the body for exit-vs-continue semantics and move the per-situation triggers fully into the reference, or vice versa.

DimensionReasoningScore

Conciseness

The body is dense and operational — no tutorial prose or explanations of concepts Claude already knows, and every section carries task-specific information. However, boundaries are repeated across sections: the mWeb/Desktop decline with `EXIT 4` appears in Scope, triage gates, and 'Out of scope'; the HybridApp gate is stated in the intro, Scope, and 'Out of scope'; the exit-code semantics are duplicated between the exit-codes table and references/error-handling.md. This fits anchor 4 (efficient with minor instances that could be trimmed) better than anchor 3, since the redundancy is reinforcement of constraints rather than unnecessary explanation.

4 / 5

Actionability

Nearly every operation is given as a copy-paste-ready command: `gh pr view <num> --json title,body`, `agent-device open "$APP_ID" --device "$DEVICE_NAME"`, `agent-device record start "$RUN_DIR/$PLATFORM/flow-$ID.mp4" --fps 24`, `test -s "$TEST_FLOW.ad" || {...}`, exact cache paths, and the fingerprint formula `sha256(precondition + json(steps) + platform)`. Concrete tables (inputs, exit codes, cost guards) cover the common cases; the few prose steps (LLM-driven flow execution) are inherently non-scriptable and are still given precise semantics (verbatim step text, append only successful actions, record chosen values in `params:`). Matches anchor 5.

5 / 5

Workflow Clarity

The sequence is explicit and ordered: numbered triage gates 'run in order, before any device work' → steps parsing → cache check → shared setup → Phase 1 → Phase 2 → manifest → handoff, with per-flow status semantics. Validation checkpoints are present throughout (Phase 1 step 6 sanity-checks the script with `test -s`, step 4 verifies final state via `agent-device is exists`, Phase 2 step 7 verifies the MP4 with `test -s` + `file`), and there are feedback loops and recovery paths (retry Phase 2 once on a 0-byte recording per references/error-handling.md, mark-and-continue per flow, exit codes 5/6 for total failures). This satisfies the batch-operation requirement — the destructive/batch cap of 3 does not apply because validation is explicit — and matches anchor 5.

5 / 5

Progressive Disclosure

The body is an orchestration overview with well-signaled, one-level-deep references, all of which exist: [`references/steps-parsing.md`], [`references/manifest-schema.md`], and [`references/error-handling.md`], each referenced from its matching section ('Steps parsing', 'Manifest schema', 'Error handling') and each self-contained with no nested references. Detail material (parsing heuristics, schema field semantics, the error matrix) is correctly split out while orchestration stays in the main file. This matches anchor 5 (clear overview, well-signaled one-level references, appropriate split).

5 / 5

Total

19

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: concrete domain and actions, explicit 'Use when...' trigger phrases that mirror real user requests, third-person voice, and a clear scope boundary against mWeb/Desktop. The only minor gap is that it describes one core capability (record evidence) rather than enumerating several distinct actions, leaving specificity just short of full marks.

DimensionReasoningScore

Specificity

Names the domain ("iOS/Android native MP4 evidence" from "an Expensify GitHub PR or issue") and concrete actions (record MP4 evidence, produce screenshots/videos, declines mWeb and Desktop), but coverage centers on a single record-and-produce action rather than a list of multiple distinct capabilities. It clearly sits above anchor 3 (only 1-2 concrete actions, not comprehensive) because it also specifies the artifact types, the source kinds, and the boundary behavior, but below anchor 5, which expects several distinct concrete actions.

4 / 5

Completeness

Explicitly answers what ("Records iOS/Android native MP4 evidence for test/repro flows extracted from an Expensify GitHub PR or issue") and when ("Use when the user asks to 'record the flow for PR #X'...") with concrete trigger phrases, exactly matching anchor 5. Not anchor 4 because the 'when' clause is fully explicit rather than merely present.

5 / 5

Trigger Term Quality

Quotes three natural user phrasings — "record the flow for PR #X", "capture mobile evidence for issue #Y", "produce screenshots/videos for <PR or issue URL>" — covering synonyms for both the verb (record/capture/produce) and the artifact (MP4, screenshots, videos), plus the source terms (PR, issue, mobile). It matches anchor 5's comprehensive natural-term coverage; no common phrasing a user would naturally say for this task is missing.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (Expensify PR/issue native-mobile evidence capture) with distinct trigger phrasing, and explicitly fences off adjacent territory ("Mobile-native only - declines mWeb and Desktop"), minimizing conflict with browser-driver skills. This matches anchor 5; anchor 4's 'minor overlap risk with closely related skills' does not apply given the explicit boundary.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

relative_links

Relative link issues: 5 suspicious

Warning

Total

14

/

16

Passed

Repository
Expensify/App
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.