CtrlK
BlogDocsLog inGet started
Tessl Logo

web-pentest

Authorized web pentest: recon, proof-based exploits, report.

57

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/security/web-pentest/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a strong, operationally dense skill body: phased workflow, hard guardrails, proof-based verdict levels, and genuine validation/stop checkpoints that are exemplary for a destructive-operation skill. The principal defects are broken bundle references — the whole templates/ directory and approval.py are cited but absent — and modest redundancy in the redaction guidance.

Suggestions

Ship the missing templates/authorization.md, templates/pentest-report.md, and templates/exploitation-queue.json files referenced in Phases 0, 3, and 5, or inline their essential contents so the workflow is self-contained.

Consolidate the credential-redaction rules stated in guardrail 6 and Phase 4 into a single authoritative location to save tokens and avoid drift.

Resolve or remove the reference to the approval.py system, which does not exist in the bundle, so destructive-payload approval has a concrete mechanism.

DimensionReasoningScore

Conciseness

The body is dense and operational with no padding explaining concepts Claude already knows; every phase carries concrete instruction. Not 5 because there is minor redundancy (the credential-redaction rule appears in both guardrail 6 and Phase 4) and environment-specific Hermes configuration detail that could be trimmed or moved to a reference.

4 / 5

Actionability

Concrete executable commands (nmap, whatweb, curl, engagement-dir setup), a verbatim authorization prompt, and specific witness payloads (' AND 1=1--, the SVG marker) cover the common cases. Not 5 because templates/authorization.md, templates/pentest-report.md, templates/exploitation-queue.json, and the approval.py system are referenced but do not exist in the bundle, leaving the queue schema and report format underspecified.

4 / 5

Workflow Clarity

Phases 0-5 are clearly sequenced with explicit validation checkpoints (authorization gate, per-request scope check, pre-send checks, bypass-exhaustion before false-positive classification) and error-recovery feedback loops plus a 'When to Stop' section. The destructive-operation cap does not apply because validation and approval gates are pervasive.

5 / 5

Progressive Disclosure

The body is well-structured with one-level-deep, clearly signaled references and a Further Reading index, but scored against the actual bundle: 3 of the 8 referenced paths (the entire templates/ set) plus approval.py are missing, so a third of the navigation targets are dead. This drags it below the 'mostly clear' anchor 4.

3 / 5

Total

16

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names a distinct niche with the full recon-to-report action pipeline, and importantly signals the authorization boundary. Its main weakness is the missing 'Use when...' trigger clause and sparse synonym coverage ('penetration testing', 'vulnerability'), which limit discoverability and completeness.

Suggestions

Add an explicit trigger clause, e.g. "Use when the user asks to pentest, penetration-test, or find vulnerabilities in a web application they own or are authorized to test."

Include common synonyms such as "penetration testing", "vulnerability assessment", and "security testing" so natural phrasings match the skill.

Slightly expand the action list to mention the report's proof-based nature, e.g. "...and produces a CVSS-scored findings report with reproducible evidence."

DimensionReasoningScore

Specificity

"Authorized web pentest: recon, proof-based exploits, report" lists several concrete actions covering the full pipeline (reconnaissance, exploitation, reporting) plus an authorization qualifier. Not 5 because the actions are compressed and 'report' is generic; not 3 because it names more than 1-2 actions with broad coverage of the domain.

4 / 5

Completeness

The 'what' is clear (authorized web pentest with recon/exploit/report phases), but there is no 'Use when...' clause or equivalent explicit trigger guidance, which caps completeness at 3 per the judging guidelines.

3 / 5

Trigger Term Quality

"web pentest" and "exploits" are natural user terms, but common variations like "penetration testing", "vulnerability assessment", and "security testing" are absent. This matches the anchor for some relevant keywords with missing variations rather than good coverage.

3 / 5

Distinctiveness Conflict Risk

"web pentest" carves out a clear niche (active web exploitation) that is mostly distinct from passive code-audit skills. Not 5 because it could mildly overlap with general security-review skills given no distinguishing trigger phrases beyond 'pentest'.

4 / 5

Total

14

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.