CtrlK
BlogDocsLog inGet started
Tessl Logo

benchmark

Benchmark mode marker — engagement objective is flag capture. Generic engagement rules apply unchanged.

58

Quality

67%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./packages/decepticon/decepticon/skills/benchmark/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a tight, highly actionable operational skill with executable flag-sweep code, a clear short-circuit workflow gated on a verified flag, and a well-structured routing table. The only real slack is some justification prose around the routing table that could be trimmed.

DimensionReasoningScore

Conciseness

The body is operational and assumes Claude's knowledge (no basic CTF/RCE explanations), but the routing-table preamble and surrounding justification prose could be tightened without losing the actionable content.

2 / 3

Actionability

Provides fully executable bash for the batched flag-path sweep and grep, with explicit guidance to replace the flagged <TARGET>/<RCE_SINK> placeholders, plus concrete numbered short-circuit steps and a tag-to-skill routing table.

3 / 3

Workflow Clarity

Sequences are clear with explicit gates: 'after RCE confirmed' precondition, 'verified flag or flag-equivalent credential' before short-circuiting, and an ordered 'flag-path first, credential second' sweep with a broad find fallback for non-standard paths.

3 / 3

Progressive Disclosure

The skill is a well-organized overview that defers detail to clearly signaled, one-level-deep peer-skill references (e.g., /skills/standard/exploit/web/<X>/SKILL.md, decepticon.md) with no nested reference chains, and a 'What this skill is NOT' section aids navigation.

3 / 3

Total

11

/

12

Passed

Description

50%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly identifies the benchmark/flag-capture niche and objective but stops short of listing concrete capabilities or giving an explicit 'Use when' trigger clause. It is adequate but generic enough to leave the activation condition implicit.

Suggestions

Add an explicit 'Use when ...' trigger clause (e.g., 'Use when running benchmark/CTF challenges or capturing flags') to clearly signal when this skill applies and lift completeness and trigger-term coverage.

List concrete actions the skill performs instead of only declaring the objective, e.g., 'routes vulnerability tags to exploit sub-skills, runs flag-path sweeps, and short-circuits to the final flag on capture.'

Include common trigger variations users would naturally say ('CTF', 'challenge', 'capture the flag') alongside 'benchmark' and 'flag capture' to broaden trigger-term coverage.

DimensionReasoningScore

Specificity

Names the domain ('benchmark mode marker', 'engagement objective is flag capture') but does not list concrete capabilities, and 'Generic engagement rules apply unchanged' is abstract rather than action-oriented.

2 / 3

Completeness

The 'what' is stated (flag capture engagement objective), but there is no explicit 'Use when...' clause, so the 'when' is only implied rather than explicit.

2 / 3

Trigger Term Quality

'benchmark' and 'flag capture' are relevant keywords a user might say, but common variations like 'CTF', 'challenge', or 'capture the flag' are absent, so coverage is incomplete.

2 / 3

Distinctiveness Conflict Risk

The benchmark/flag-capture niche is fairly distinct, but the statement that generic engagement rules apply unchanged signals reliance on and overlap with the broader engagement skill.

2 / 3

Total

8

/

12

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
PurpleAILAB/Decepticon
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.