CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-device-evidence

Records iOS/Android native MP4 evidence for test/repro flows extracted from an Expensify GitHub PR or issue. Use when the user asks to "record the flow for PR

76

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

High

Do not use without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured autonomous workflow with executable commands, pervasive validation checkpoints, and clean one-level-deep references. The only weakness is minor redundancy between the Scope and 'Out of scope' sections.

Suggestions

Consolidate the 'Out of scope (do not do these)' section with the earlier 'Out of scope' subsection in Scope to remove the repeated mWeb/Desktop and HybridApp-gate declinations.

Consider extracting the per-phase 'agent-device open' invocation into a single named snippet referenced by Phase 1 and Phase 2 to avoid restating the identical command.

DimensionReasoningScore

Conciseness

Dense and operational with no conceptual padding, but the 'Out of scope (do not do these)' section repeats material already stated in Scope and the open command is restated across phases; minor trimming possible.

4 / 5

Actionability

Executable, copy-paste-ready commands throughout - agent-device open/close/record/replay, mkdir, test -s, file - with real flags and paths covering the common phase cases.

5 / 5

Workflow Clarity

Clearly sequenced triage gates, shared setup, and two phases with explicit validation checkpoints (test -s guards, verify-final-state, non-empty-script sanity check) and per-flow failure feedback loops plus an exit-code table.

5 / 5

Progressive Disclosure

Lean overview with three well-signaled, one-level-deep references (steps-parsing.md, manifest-schema.md, error-handling.md), each linked inline at the relevant step; all referenced files exist and hold the expected detail.

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A high-quality description: concrete actions, natural trigger phrases, explicit what-and-when guidance, and a sharp scope boundary that reduces conflict risk. All in third person with no padding.

DimensionReasoningScore

Specificity

Lists multiple concrete actions - 'Records iOS/Android native MP4 evidence', 'produce screenshots/videos' - tied to a specific source (PR or issue), matching the comprehensive-coverage anchor.

5 / 5

Completeness

Explicitly answers both 'what' (records native MP4 evidence from PR/issue flows) and 'when' with concrete trigger phrases, plus a scope-declination clause.

5 / 5

Trigger Term Quality

Includes natural user phrasings - 'record the flow for PR #X', 'capture mobile evidence for issue #Y', 'produce screenshots/videos for <URL>' - covering synonyms across the same intent.

5 / 5

Distinctiveness Conflict Risk

A clearly bounded niche (Expensify mobile-native evidence from GitHub PR/issue) with explicit declination of mWeb/Desktop, minimizing overlap with other skills.

5 / 5

Total

20

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

relative_links

Relative link issues: 5 suspicious

Warning

Total

14

/

16

Passed

Repository
Expensify/App
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.