CtrlK
BlogDocsLog inGet started
Tessl Logo

browser-testing

Drive real browsers via Chrome DevTools MCP: navigate pages, capture snapshots, run responsive checks, and collect console/perf traces. Use when the user mentions: 'validate UI change in Chrome', 'capture a screenshot', 'run responsive checks', or 'collect console logs'. Trigger terms: browser testing, DevTools, console logs, screenshot, responsive testing

75

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

88%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An efficient, highly actionable skill body: concrete tool parameter shapes, a screenshot budget constraint, a sequenced workflow with error-recovery loops, and hard validation gates. The two minor gaps are the absence of a complete evaluate_script example and an external config reference that cannot be verified in the bundle.

Suggestions

Include one complete, copy-paste-ready evaluate_script call (e.g. the exact {function: '() => !!document.querySelector(...)'} payload) so the arrow-function-string convention is unambiguous.

Clarify where insightSetId comes from in the performance trace flow (e.g. the return value of performance_start_trace) to make the perf-trace steps fully executable.

Verify the .opencastle/stack/testing-config.md path exists in the target project (or mark it as project-provided) so the lone external reference is not a dead pointer.

DimensionReasoningScore

Conciseness

The body is lean and operational throughout — "Screenshots are expensive. **MAX 3 per session**, reserved for failures" — with zero background on what Chrome DevTools or MCP is, and no padded sections. Every line teaches something Claude would not already know (uid-not-selector semantics, screenshot budget, wait_for failure diagnosis), matching the 'every token earns its place' anchor; it is not 4 because no section reads as over-explanation.

5 / 5

Actionability

Tool calls come with concrete parameter shapes ("`click` and `type` take a `uid` from a prior snapshot, not a CSS selector", "`evaluate_script` — `{ function: '() => ...' }` (an arrow function *string*)"), plus named assertion patterns. It falls short of the fully copy-paste-ready anchor: there is no complete example `evaluate_script` invocation, and the perf-trace flow references `insightSetId` without stating where it comes from — minor gaps rather than vague guidance.

4 / 5

Workflow Clarity

The workflow is a clearly sequenced chain ("Navigate → `wait_for` anchor text → assert via `evaluate_script` → ... → `list_console_messages`") with an explicit error-recovery feedback loop ("any error: fix source, rebuild, reload, restart from navigate") and hard validation gates ("Every test must pass before writing the updated `result.json`. Do not stop on partial green"), plus a troubleshooting heuristic for wait_for timeouts. This matches the anchor requiring explicit validation steps and feedback loops; no destructive/batch cap applies.

5 / 5

Progressive Disclosure

The skill is under 50 lines with no bundle files and well-organized sections, which would qualify for the top anchor — but the single referenced path, ".opencastle/stack/testing-config.md", does not exist in this tree, leaving the one external pointer unverifiable. That is a minor organization gap versus the 'well-signaled one-level-deep references' anchor rather than a structural problem.

4 / 5

Total

18

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete capability list, explicit quoted user triggers, and a clear tool-specific niche in third-person voice. The only weakness is that a couple of the broad trigger terms (screenshot, console logs) could fire in non-browser contexts.

DimensionReasoningScore

Specificity

The description lists four concrete, specific actions — "navigate pages, capture snapshots, run responsive checks, and collect console/perf traces" — tied to a named mechanism (Chrome DevTools MCP). This matches the anchor for multiple specific concrete actions with comprehensive coverage; it is above the 'several specific actions, minor gaps' anchor because the capability list covers the tool's core surface.

5 / 5

Completeness

Both questions are explicitly answered: the 'what' (drive real browsers via Chrome DevTools MCP with the four listed capabilities) and the 'when' ("Use when the user mentions: 'validate UI change in Chrome', 'capture a screenshot', ..." with concrete trigger phrases), structurally matching the top anchor's example. It is not 4 because the 'when' clause is fully explicit with quoted user phrases rather than merely adequate.

5 / 5

Trigger Term Quality

It provides natural user-voice phrases ("validate UI change in Chrome", "capture a screenshot", "collect console logs") plus an explicit synonym-inclusive trigger list (browser testing, DevTools, console logs, screenshot, responsive testing), matching the comprehensive-coverage-with-synonyms anchor rather than the 'a few natural terms missing' anchor.

5 / 5

Distinctiveness Conflict Risk

The niche is clear and tool-specific ("Drive real browsers via Chrome DevTools MCP"), but generic trigger terms like "screenshot" and "console logs" create minor overlap risk with screenshot-capture or logging-related skills. This fits the 'mostly distinct; minor overlap risk with closely related skills' anchor better than the minimal-conflict top anchor.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
monkilabs/opencastle
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.