CtrlK
BlogDocsLog inGet started
Tessl Logo

exploratory-test

Explore a running Positron instance as a real user to find genuine problems in a change you just made. Use when asked to exploratorily test, QA, manually test, or poke at a branch, PR, or feature through the real UI. This is discovery testing against the live app to find bugs, NOT writing automated tests; use author-e2e-tests or author-vitest-tests for that. Worth its cost for a user-visible behavior change, not for a refactor or a typo fix. Only runs when a person invokes it explicitly.

76

Quality

95%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, fully executable orchestration skill with strong error-recovery checkpoints and excellent token efficiency — it reads as a dispatch plan rather than a manual. The single weakness is that all referenced bundle files are absent from the provided skill directory, leaving the disclosed content unverifiable.

Suggestions

Ship the referenced files (explorer.md and the renderer/ scripts: render.mjs, finish.mjs, publish.sh) alongside SKILL.md, or list them in the frontmatter, so the one-level-deep disclosure can actually be followed.

Use the full path form consistently in the render command — "node <render.mjs> <report.md> ..." should read "node <base>/renderer/render.mjs ..." as the neighboring finish.mjs and publish.sh commands do.

DimensionReasoningScore

Conciseness

Every sentence carries operational information the model could not infer (fork cost rationale, why the brief must be self-contained, why re-render needs the harness-side totals); nothing explains concepts Claude already knows. Not 4: there are no over-explanations to trim.

5 / 5

Actionability

Fully executable guidance: exact subagent configuration ("subagent_type: \"general-purpose\"" and "model: \"opus\""), copy-paste-ready commands with full flag lists ("node <base>/renderer/finish.mjs prompt <run dir> --repo <checkout> --base <base sha> --head <head sha>"), and even the exact wording to ask the user. Not 4: no gaps in the command coverage.

5 / 5

Workflow Clarity

The explore → verify → apply → render → publish sequence is explicit with validation checkpoints and error-recovery feedback loops at every risky step: "If it prints no findings, skip to the render", "If it says there are no AWS credentials, relay its sign-in hint", "If it stops on a screenshot it could not paint a credential out of... leave the report unpublished", and "Publish only on a yes". Not 4: no validation gap remains.

5 / 5

Progressive Disclosure

Good structure: a lean orchestration body pointing one level deep to clearly signaled external files ("<base>/explorer.md: tell the agent to read it in full before anything else... none of that is in this file"). Not 5: the referenced files (explorer.md, renderer/render.mjs, renderer/finish.mjs, renderer/publish.sh) are not present in the bundle as provided, so the split cannot be verified and navigation depends on unshipped assets.

4 / 5

Total

19

/

20

Passed

Description

95%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: it states what and when concretely, uses natural trigger synonyms, and actively disambiguates from neighboring test-authoring skills. The only minor gap is that the action list is a single core capability rather than multiple distinct actions.

DimensionReasoningScore

Specificity

Concrete actions are named — "Explore a running Positron instance as a real user to find genuine problems" and "discovery testing against the live app to find bugs" — but the action list is narrower than the 5-anchor's comprehensive coverage. It is well above the 3-anchor's 1-2 actions because the scope, target, and outcome are all specific.

4 / 5

Completeness

Explicitly answers both — what ("Explore a running Positron instance as a real user to find genuine problems in a change") and when ("Use when asked to exploratorily test, QA, manually test, or poke at a branch, PR, or feature") — with concrete trigger phrases and an explicit anti-trigger ("NOT writing automated tests").

5 / 5

Trigger Term Quality

"Use when asked to exploratorily test, QA, manually test, or poke at a branch, PR, or feature through the real UI" covers comprehensive natural synonyms users would actually say. Not 4: no common variation is missing.

5 / 5

Distinctiveness Conflict Risk

Clear niche (exploratory testing of a live Positron instance) with explicit disambiguation from adjacent skills ("use author-e2e-tests or author-vitest-tests for that"), making wrong-skill triggering minimal. Third person voice is used throughout.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
posit-dev/positron
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.