CtrlK
BlogDocsLog inGet started
Tessl Logo

octopus-security-audit

OWASP compliance, vulnerability scanning, and adversarial red team testing — use for security reviews

57

Quality

65%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/octopus-security-audit/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, command-rich security audit body with concrete usage examples, clear mode-escalation rules, and a validated adversarial workflow. Its weaknesses are padding with concepts Claude already knows (STRIDE, OWASP category lists), non-executable guidance in the CI/CD section, and cross-references to bundle files that do not exist alongside no actual progressive disclosure into reference files.

Suggestions

Move the STRIDE table, OWASP Top 10 list, and detailed capability listings into a reference file (or trim them to one line) — Claude already knows these frameworks, so the inline tables cost tokens without adding guidance.

Ship the referenced files (skills/blocks/codex-host-adapter.md, skills/blocks/fable5-prompting.md, agents/personas/security-auditor.md) in the bundle, or remove the references, so progressive disclosure resolves to real files.

Replace the CI/CD 'check for dangerous patterns' comment list with executable detection commands (e.g., grep patterns for pull_request_target, permissions: write-all, unpinned uses:) to match the actionability of the secrets-archaeology section.

DimensionReasoningScore

Conciseness

The body is mostly efficient — commands and escalation rules are dense and useful — but includes content Claude already knows: a full STRIDE category table ('Spoofing: Can an attacker impersonate...'), a plain OWASP Top 10 category list, and a generic capabilities list. This matches 'mostly efficient but includes some unnecessary explanation or could be tightened' rather than 4, where such trimmable sections would be minor.

3 / 5

Actionability

Mostly executable guidance: concrete `orchestrate.sh` invocations, copy-paste-ready `git log -S` secrets-archaeology commands, and find/grep commands for CI/CD auditing. Gaps keep it below 5: the CI/CD section lists patterns to 'check for' as comments without detection commands, and the squeeze phases (blue/red/remediation/validation) are described at a high level with only a single invocation per workflow.

4 / 5

Workflow Clarity

Sequences are clear: quick vs. deep mode with explicit auto-escalation rules, and the squeeze workflow is a 4-phase cycle ending in 'Validation (Verify): re-tests, confirms fixes or fails' — a validation checkpoint. The Fable routing section even defines retry-once/fail-closed handling. It stops short of 5 because there are no explicit error-recovery checkpoints for the quick scan path or for a failed orchestrate.sh invocation.

4 / 5

Progressive Disclosure

The body has good section structure, but everything lives inline in one ~180-line SKILL.md with no bundle files (no references/, scripts/, or assets/ exist), while cross-references like 'skills/blocks/codex-host-adapter.md', 'skills/blocks/fable5-prompting.md', and 'agents/personas/security-auditor.md' point to files that are not part of the bundle. That is 'some structure but could be better organized' with dangling references — inline content (STRIDE table, OWASP list) that belongs in reference files.

3 / 5

Total

14

/

20

Passed

Description

67%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A terse, third-person description that states what the skill does and includes an explicit 'use for' trigger. Its main weaknesses are missing common synonyms (penetration testing, security audit) and a generic when-clause that could be more concrete. No fluff or over-claims.

Suggestions

Add natural trigger synonyms such as 'penetration testing', 'pentest', or 'security audit' so the description matches phrasings users actually say.

Make the 'when' clause more concrete, e.g. 'Use when the user asks for a security audit, vulnerability scan, or red team review of code, auth, or CI/CD changes'.

DimensionReasoningScore

Specificity

The description lists three concrete activities — 'OWASP compliance, vulnerability scanning, and adversarial red team testing' — which matches the anchor 'lists several specific actions; minor gaps in coverage'. It falls short of a 5 because 'OWASP compliance' is domain-shorthand rather than a specific action, and no detail is given on what the scanning covers.

4 / 5

Completeness

Both parts are explicit: what ('OWASP compliance, vulnerability scanning, and adversarial red team testing') and when ('use for security reviews'). The 'when' clause is present but generic — it could name concrete trigger situations (audit request, pre-deployment review) — matching 'has both what and when; when could be more explicit or specific', not the 5 anchor's concrete trigger phrases.

4 / 5

Trigger Term Quality

Relevant keywords like 'vulnerability scanning', 'red team testing', and 'security reviews' are present, but common user synonyms are missing — 'penetration testing'/'pentest', 'security audit', 'threat model', 'hardening'. This fits 'some relevant keywords but missing common variations or synonyms' rather than 4, which requires only a few natural terms missing.

3 / 5

Distinctiveness Conflict Risk

Terms like 'OWASP', 'red team testing', and 'vulnerability scanning' carve a clear security niche, but 'use for security reviews' is a broad trigger that could also fire for a general code-review skill — 'mostly distinct; minor overlap risk with closely related skills'. Not a 5 because the trigger clause lacks the distinct, narrow phrasing of the top anchor.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.