CtrlK
BlogDocsLog inGet started
Tessl Logo

pentest-agent-os

渗透测试Agent操作系统。把渗透抽象成状态空间搜索:不预设固定路径,以项目黑板 Fact 图 (upsert_project_fact)+漏洞记录沉淀认知,路径从已验证事实上涌现。覆盖全杀伤链攻击手法 (联网情报/Web/认证/服务端/源码/社工/后渗透/二进制/内网域/云/区块链/AI/无线/硬件)+0day+ 组合拳+代理自举。核心:全网搜不到洞时现场推导独属于目标的攻击链。本文件为套件索引。 Use when starting a full-chain pentest engagement or needing the skill map for this suite.

56

Quality

64%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./skills/pentest-agent-os/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an efficient, well-structured suite index that successfully routes to sibling skills, but it stops short of being mechanically actionable: skills are named without paths, the usage sequence repeats a step, and the routing decision (which skill for which surface) is only exemplified, not specified. Tightening the routing table into a decision checkpoint with resolvable references would lift it.

Suggestions

Give each sibling skill a resolvable reference (file path or link) rather than a bare name, so loading a sub-skill is mechanical instead of dependent on the harness resolving skill names.

Remove the duplicated '按攻击面加载对应 skill' step (step 4 repeats step 2) and the self-referential first table row to tighten the index.

Turn the per-surface routing into an explicit decision checkpoint — e.g. a condition→skill mapping for 'attack surface identified as X → load Y' — rather than three inline examples, and state where the full trigger text is loaded from.

DimensionReasoningScore

Conciseness

The body is lean for what it carries: a mapping table, a four-step usage sequence, and a three-line trigger quick reference, with no filler explaining concepts Claude already knows. It is not a 5 due to small redundancies — step 2 '按当前攻击面加载对应 skill' repeats almost verbatim as step 4, and the first table row ('套件索引与核心心法 | pentest-agent-os | 套件索引与核心心法') restates its own topic column.

4 / 5

Actionability

Routing guidance is partly concrete — '先加载 pentest-agent-os(本索引)', named per-surface skills with examples ('识别组件→component-vuln-intel;Web→web-attack-methods'), and triggers tied to actions ('执行 component-vuln-intel 全部命令', 'upsert_project_fact'). But sibling skills are referenced by name only with no paths or loading mechanism, and the triggers defer detail elsewhere ('完整原文在 pentest-blackboard'), leaving key execution details missing — matching 'some concrete guidance but incomplete' rather than 4.

3 / 5

Workflow Clarity

A numbered usage sequence exists and the trigger quick reference does encode validation feedback ('线索 tentative,验证后再 confirmed', '不通则写负结果 Fact'). However, the sequence duplicates itself (steps 2 and 4 both say load the skill for the attack surface), and how to decide which skill applies beyond three examples is left implicit — sequence present but checkpoints are partially implicit, fitting 3 rather than 4's 'most checkpoints present'.

3 / 5

Progressive Disclosure

The body is a well-organized index: a topic→skill→role table, explicit usage steps, and a clearly signaled pointer that full trigger text lives in `pentest-blackboard`. No bundle files exist in references/, scripts/, or assets/, so all detail is appropriately deferred to sibling skills; the gap keeping it from 5 is that sibling skills are referenced by name only, with no file paths or links to make navigation mechanical.

4 / 5

Total

14

/

20

Passed

Description

67%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description answers both what and when explicitly and stays in third person, but its trigger vocabulary is thin relative to the breadth of the suite, and the abstract state-space-search framing dilutes operational specificity. It reads as a capable suite index with room to add natural synonyms and tighten the what/when boundary against sub-skills.

Suggestions

Add natural trigger synonyms users would actually say — 'penetration testing', 'red team engagement', 'security assessment' — alongside the existing 'pentest' terms.

Replace or supplement abstract framing ('把渗透抽象成状态空间搜索', '组合拳') with one or two concrete operational actions a user can recognize.

Make the 'when' clause explicitly distinguish this index skill from the domain sub-skills (e.g. 'Use this index when beginning any engagement; load sub-skills after the attack surface is mapped').

DimensionReasoningScore

Specificity

The description lists several concrete capabilities: '以项目黑板 Fact 图 (upsert_project_fact)+漏洞记录沉淀认知', '覆盖全杀伤链攻击手法 (联网情报/Web/认证/服务端/源码/...)', '0day+组合拳+代理自举', '全网搜不到洞时现场推导独属于目标的攻击链'. It stops short of a 5 because much of the phrasing is abstract framing ('把渗透抽象成状态空间搜索', '路径从已验证事实上涌现') rather than operational actions, and it is clearly above a 3 since multiple specific actions and enumerated domains are named.

4 / 5

Completeness

Both parts are present: a clear 'what' ('渗透测试Agent操作系统...覆盖全杀伤链...本文件为套件索引') and an explicit 'when' ('Use when starting a full-chain pentest engagement or needing the skill map for this suite'). Not a 5 because the 'when' clause covers only two fairly narrow triggers and omits the most common user phrasings (e.g. when asked to test a target's security or find vulnerabilities).

4 / 5

Trigger Term Quality

Relevant keywords exist — 'pentest', 'full-chain pentest engagement', 'skill map' — but common variations a user would naturally say are missing: 'penetration testing', 'red team/red teaming', 'security assessment', 'vulnerability assessment'. It sits at 'some relevant keywords but missing common variations or synonyms', not at 4 where only a few natural terms would be missing.

3 / 5

Distinctiveness Conflict Risk

The niche is clear — an index/orchestrator for a pentest suite, explicitly flagged as '本文件为套件索引' with a trigger ('needing the skill map for this suite') that distinguishes it from the domain sub-skills. It is not a 5 because the description's scope language ('覆盖全杀伤链攻击手法...0day+组合拳') overlaps heavily with the enumerated sibling skills, creating some risk of triggering when a sub-skill is wanted.

4 / 5

Total

15

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

Total

14

/

16

Passed

Repository
AIPentest/CyberStrikeAI
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.