CtrlK
BlogDocsLog inGet started
Tessl Logo

web-pentest

Perform black-box / grey-box web application penetration testing on an authorized target — auth bypass, IDOR, session handling, business-logic flaws, parameter tampering, Burp Suite / OWASP ZAP workflows. Use when the user mentions 'web pentest,' 'web application penetration test,' 'pentesting,' 'bug bounty,' 'Burp Suite,' 'ZAP,' 'OWASP testing,' 'authentication testing,' 'session testing,' 'authorization testing,' 'business logic testing,' 'web vulnerability testing,' or has explicit authorization to test a live web application.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable methodology document: concrete payloads, commands, expected-response checkpoints, and responsible boundaries. Its weaknesses are structural — the whole methodology lives inline in SKILL.md with no progressive disclosure into reference files — and occasional prose that could be tightened.

Suggestions

Move the payload banks (SQLi, XSS, deserialization markers) and the detailed tool guide into references/ files (e.g., references/payloads.md, references/tools.md), leaving SKILL.md as a lean phase-by-phase overview that points to them — this is the main progressive-disclosure gap.

Tighten the Tooling section to one line per tool (name — purpose) and trim flavor commentary (e.g., "look for the one endpoint that didn't get the memo") to reduce token cost without losing the heuristics.

Add explicit feedback loops per phase, e.g., "if a 200 is observed on a cross-tenant request, immediately re-verify with a fresh account pair before reporting" and "after any Intruder run, confirm no test data was modified" — to reach full workflow-clarity validation coverage.

DimensionReasoningScore

Conciseness

The body is a dense, actionable checklist with almost no re-teaching of concepts Claude already knows, but there are minor trimmable instances — tool prose like "dalfox / XSStrike — XSS scanners (high false-positive; use as a starting point)" and flavor commentary such as "look for the one endpoint that didn't get the memo." This fits the 4 anchor (efficient with minor over-explanation) rather than 5, where every token would earn its place.

4 / 5

Actionability

Fully executable guidance throughout: concrete commands ("curl -I -H \"Origin: https://evil.com\"", "testssl.sh https://target"), copy-paste payloads ("' OR 1=1 --", "'; SELECT pg_sleep(5)--", "{{7*7}}"), serialized-format markers, sqlmap risk/level flags, and a step-by-step IDOR test procedure with expected response codes — matching the 5 anchor's copy-paste-ready coverage of common cases.

5 / 5

Workflow Clarity

Nine clearly sequenced WSTG phases with an upfront authorization gate, per-test expected outcomes ("200 (bad), 403 (good)"), explicit verification probes ("replay the cookie post-logout"), and a report template with a Verification field. It falls short of the 5 anchor because error-recovery feedback loops are mostly implicit — only the active-compromise case has an explicit stop-and-notify loop.

4 / 5

Progressive Disclosure

Sections are well-organized with clear headers, but the entire ~200-line methodology is inlined in SKILL.md with no bundle files at all — payload banks, the tool guide, and the report template are natural candidates for references/ files. This matches the 3 anchor (some structure, content that should be separate is inline), not 4 (nothing is actually split out or signaled as external).

3 / 5

Total

16

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete capability list, explicit and extensive trigger guidance, third-person imperative voice, and a clearly defined niche. The only weakness is mild trigger overlap with adjacent security-audit skills on the broader terms like "pentesting" and "web vulnerability testing."

DimensionReasoningScore

Specificity

The description names the domain and six concrete actions — "auth bypass, IDOR, session handling, business-logic flaws, parameter tampering, Burp Suite / OWASP ZAP workflows" — giving comprehensive, specific coverage that matches the top anchor; there are no significant gaps that would drop it to 4.

5 / 5

Completeness

It explicitly answers both questions: the "what" (perform black-box/grey-box web pentest with the listed capabilities) and a literal "Use when the user mentions..." clause with concrete trigger phrases, exactly the shape of the 5 anchor.

5 / 5

Trigger Term Quality

Twelve-plus natural trigger phrases including synonyms and tool names ("web pentest," "pentesting," "bug bounty," "Burp Suite," "ZAP," "OWASP testing," "authorization testing") plus a state-based trigger ("has explicit authorization to test a live web application") — comprehensive synonym coverage matching the 5 anchor rather than the 4 anchor's "a few natural terms missing."

5 / 5

Distinctiveness Conflict Risk

The niche (live web application pentesting) and tool-specific triggers are distinct, but broad terms like "pentesting," "web vulnerability testing," and "OWASP testing" overlap with closely related security skills (the body itself references companion skills recon and owasp-audit), fitting the 4 anchor "mostly distinct; minor overlap risk" rather than the 5 anchor's minimal conflict risk.

4 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

15

/

16

Passed

Repository
briiirussell/cybersecurity-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.