Use when a GitHub issue, PR, or bug report links a screen-recording video (github.com/user-attachments/assets/... .mov/.mp4/.webm) and you need to know what it shows. You cannot view video in context — this extracts still frames to a temp dir so you can Read them as images. Trigger whenever a task references a video/screen recording/attachment demonstrating a bug, or when text says "see the recording"/"linked video".
71
86%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
Issues often attach a screen recording as the clearest record of a bug. You cannot watch video in context, but you can Read image frames. This skill extracts frames from a video (a GitHub attachment URL or local file) into a temp workspace, then reads them.
Core principle: Do it autonomously and without polluting the repo, then interpret the frames against the issue's written words — never from the frames alone. Sparse frames drop the causal step (a click and the UI it triggers land between samples), so a frames-only reading is a guess. The written repro steps are ground truth for what action was taken; the frames show the result.
Copy this checklist and track progress:
- [ ] 1. Read the issue text FIRST (title, body, repro steps, expected vs actual)
- [ ] 2. Extract frames with the script (2 fps whole-clip pass)
- [ ] 3. Read the frames; map each to a repro step
- [ ] 4. If a transition is ambiguous, zoom-extract that window (10 fps)
- [ ] 5. Reconcile: does my reading match the written repro? If not, resolve before concluding
- [ ] 6. Report frame-by-frame; clean up the temp workspaceStep 1 — Read the issue text first. Get the title, body, numbered repro steps, and expected-vs-actual. This tells you what page and what action to expect. Do NOT skip this — it is what keeps step 5 honest.
Step 2 — Extract frames. Run the bundled script (it installs ffmpeg if missing on macOS, downloads to a mktemp -d dir OUTSIDE the repo, prompts for nothing):
.claude/skills/watching-issue-video-recordings/scripts/extract-frames.sh <url-or-path> [fps] [start] [duration]For a GitHub attachment, pass the github.com/user-attachments/assets/<uuid> URL verbatim. Default 2 fps is right for a first pass. The script prints WORKSPACE=, META:, FRAMES=, then one absolute frame path per line.
Step 3 — Read the frames. Read the printed frame paths as images (batch several Read calls). Note the URL bar, visible controls, form field contents, cursor position, and any modal.
Step 4 — Zoom on ambiguity. If you can't tell which control was clicked or what triggered a modal, re-run over just that window at high fps — e.g. ... 10 1.5 1.5 samples 10 fps from 1.5s for 1.5s. The click and its effect are usually 1–2 frames apart; 2 fps can merge them.
Step 5 — Reconcile with the text (the load-bearing step). State what you think happened, then check it against the repro steps from step 1. If the recording seems to show something the text doesn't describe (e.g. "the button does nothing" when the issue says "clicking X shows the wrong modal"), you are probably misreading a merged transition — zoom in (step 4) before concluding. A confident frames-only narrative that contradicts the written steps is the most common failure of this task.
Step 6 — Report and clean up. Give a concise frame-by-frame account tied to timestamps/steps. Then remove the workspace by its literal path: rm -rf <the WORKSPACE path the script printed>. Do not glob-delete other temp dirs — unrelated issue-video.* dirs from other sessions may exist; leave any you did not create.
mktemp -d under the system temp dir. If you extract manually, do the same — a downloaded .mov or a frames/ dir committed to the repo is a defect.curl -L follows the redirect to blob storage.git status --porcelain before reporting done.| Need | Command |
|---|---|
| First pass, whole clip | extract-frames.sh <url> (2 fps default) |
| Zoom a transition (2s–3.5s) | extract-frames.sh <url> 10 2 1.5 |
| Local file instead of URL | extract-frames.sh /path/to/clip.mov |
| Clean up | rm -rf <WORKSPACE printed by the script> |
On issue 6061, the attached recording was the key to the diagnosis: a frame showed the Notes field already contained text (asdfasdfasdfasdf) before "Edit" was clicked — the precondition (hasChanged) that made the navigation guard fire. A baseline agent working from a sparse 1 fps pass, without reading the issue text, concluded the bug was "'Yes, Cancel' does nothing" — the opposite of the real Edit→wrong-modal bug — because the click→modal transition fell between samples. Reading the text first and zooming on the transition (steps 1, 4, 5) prevents that.
b61c72b
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.