Explore a running Positron instance as a real user to find genuine problems in a change you just made. Use when asked to exploratorily test, QA, manually test, or poke at a branch, PR, or feature through the real UI. This is discovery testing against the live app to find bugs, NOT writing automated tests; use author-e2e-tests or author-vitest-tests for that. Worth its cost for a user-visible behavior change, not for a refactor or a typo fix. Only runs when a person invokes it explicitly.
76
95%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Passed
No findings from the security scan
Explore a running Positron instance as a real user and find genuine problems in what you were pointed at: the change's diff plus any other features and functions in its blast radius, or the feature named in the request. Not the rest of Positron.
Spawn one fresh agent with subagent_type: "general-purpose" and
model: "opus". Do not fork: a fork costs twice the calls for fewer findings,
because it re-sends your whole conversation on every turn. Sonnet is only for a
narrow re-test of one known scenario; it is not good enough for discovery.
The brief is the only context the agent has, so make it self-contained: the
checkout path, the branch, the base and head SHAs and the git diff that shows
the change between them, what the change is meant to
do as a user would describe it, and the blast radius you are nervous about.
State intent and risk; do not state what you expect to work.
If the person gave a time limit ("spend 20 minutes on it"), put it in the brief
as minutes to explore, with how to keep to it: run date at the start and
between scenarios, and when the time is up, stop exploring, list what was not
reached under Not run, and write up. Nothing enforces it locally, as it does in
CI, so say it plainly.
Resolve two absolute paths from this skill's base directory and put both in the brief. The branch under test may predate them, so the agent cannot find them from there.
<base>/explorer.md: tell the agent to read it in full before anything
else. It is how to drive the app and the report and ledger it must write;
none of that is in this file.<base>/renderer/render.mjs: the report renderer.Running it in a subagent keeps screenshots, snapshots, and dead ends out of the session you are working in.
When the agent finishes, have a second agent check its findings, as CI does.
With the base and head SHAs from the brief, run:
node <base>/renderer/finish.mjs prompt <run dir> --repo <checkout> --base <base sha> --head <head sha>.
If it prints no findings, skip to the render. Otherwise it prints the path of
a prompt file. Spawn a fresh agent with subagent_type: "general-purpose" and
model: "sonnet", tell it to read that file and do what it says, and save its
reply exactly as returned to <run dir>/verify-reply.md. Then run
node <base>/renderer/finish.mjs apply <run dir> <run dir>/verify-reply.md.
The verdicts are advisory: do not edit them or drop a finding over them.
Then put both runs on the report's Run tile, as CI does. Each agent's
completion notice carries duration_ms and tool_uses; re-render with them,
leaving out the --verify-* flags when there was nothing to verify:
node <render.mjs> <report.md> --model <model id> --duration-ms <duration_ms> --turns <tool_uses> --verify-model <model id> --verify-duration-ms <duration_ms> --verify-turns <tool_uses>.
The agents cannot do this themselves, because they do not see their own totals.
Give the user the result, the findings table, and the index.html path. Then
ask, in these words:
Publish this report to share it? (Anyone with the link can view it. Keys found in screenshots are painted over.)
Publish only on a yes:
bash <base>/renderer/publish.sh <run dir>. It prints the report URL; give it
to the user. If it says there are no AWS credentials, relay its sign-in hint.
If it stops on a screenshot it could not paint a credential out of, name the
shot and leave the report unpublished.
b0258bc
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.