CtrlK
BlogDocsLog inGet started
Tessl Logo

web-pentest

Authorized web application penetration testing — reconnaissance, vulnerability analysis, proof-based exploitation, and professional reporting. Adapts Shannon's "No Exploit, No Report" methodology with hard guardrails for scope, authorization, and aux-client leakage. Active testing against running applications you own or have written authorization to test.

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable phased pentesting skill with strong validation gates and clear progressive disclosure of reference material. The notable weakness is missing template/queue files that the workflow repeatedly depends on, plus minor over-explanation in a few sections.

Suggestions

Add the missing bundle files referenced by the workflow: templates/authorization.md, templates/pentest-report.md, and templates/exploitation-queue.json, or remove the references if those artifacts are produced inline.

Trim explanation Claude already knows — e.g., the -T3 vs -T4 IDS rationale and the extended Hermes aux-client mitigation aside — or move them into a references file.

Consider moving the long guardrails block into references/scope-enforcement.md and keeping a short inlined checklist, to reduce token load on routine invocations.

DimensionReasoningScore

Conciseness

Mostly efficient and dense with high-value phased content, but includes some explanation Claude does not need (e.g., justifying -T3 over -T4/-T5 to avoid IDS, and the Hermes aux-client leakage aside), and the guardrails section could be tightened.

3 / 5

Actionability

Fully executable throughout: concrete nmap/whatweb/curl commands, exact witness payloads (' AND 1=1--', the benign XSS marker), delegate_task subagent table with reference columns, queue field schemas, and per-phase output file paths — copy-paste ready and covering the common cases.

5 / 5

Workflow Clarity

Phases 0–5 are clearly sequenced with explicit validation checkpoints (authorization gate, pre-send scope/auth/payload checks, bypass exhaustion before false-positive, 'When to Stop' conditions) and feedback loops, so the destructive/batch cap at 3 does not apply.

5 / 5

Progressive Disclosure

Good one-level-deep structure with a 'Further Reading' index and real references/ and scripts/ files, but three paths the body relies on (templates/authorization.md, templates/pentest-report.md, templates/exploitation-queue.json) are missing from the bundle, breaking navigation to core templates.

4 / 5

Total

17

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that clearly conveys both purpose and trigger conditions with a distinctive niche. The main weakness is second-person voice ('you own / you have written authorization') where third-person is expected, which costs a specificity point.

Suggestions

Rewrite in third person to avoid the voice penalty: e.g. 'Performs authorized web application penetration testing against applications the operator owns or has written authorization to test.'

Surface the most natural trigger terms ('pentest', 'security test', 'find vulnerabilities') directly in the description text rather than only in the triggers list.

DimensionReasoningScore

Specificity

Names the domain and several concrete actions ('reconnaissance, vulnerability analysis, proof-based exploitation, and professional reporting') with guardrails, but second-person phrasing ('you own', 'you have written authorization') triggers the voice penalty, reducing the score by one from a strong 4.

3 / 5

Completeness

Explicitly answers both 'what' (recon, vuln analysis, proof-based exploitation, reporting) and 'when' ('Active testing against running applications you own or have written authorization to test') with concrete trigger framing.

5 / 5

Trigger Term Quality

Strong keyword coverage via the frontmatter triggers ('pentest', 'penetration test', 'security test', 'find vulns', 'OWASP test') plus the description's 'penetration testing' and 'vulnerability', though a few natural synonyms are absent from the description text itself.

4 / 5

Distinctiveness Conflict Risk

A clear niche (authorized web app pentesting with proof-based exploitation and hard scope/authorization guardrails) with distinct triggers and minimal overlap with other skills.

5 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.