CtrlK
BlogDocsLog inGet started
Tessl Logo

testing-enemy-arrogance

Use when probing enemy leadership's judgment quality and decision-making state. Based on Zhong's grain-request test against Wu: make a resource request, observe whether wise counsel is heeded or ignored, and interpret the response to assess arrogance and internal divisions for attack timing.

70

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

85%

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is well-organized and lean, with a clear sequenced workflow and a built-in verification section. Its main weakness is actionability: several steps are phrased as abstract directives rather than concrete, executable procedures.

Suggestions

Make each step more operational—specify what exactly to request, how to observe the decision process, and what signal to record—rather than verbs like 'observe' and 'monitor'.

Add a short concrete checklist of evidence to capture (who advised caution, how the leader responded, what was granted/denied) so the procedure is copy-paste actionable.

DimensionReasoningScore

Conciseness

The body is lean and avoids explaining background concepts Claude already knows; each section adds only the procedural detail needed, matching the 'lean and efficient' anchor.

3 / 3

Actionability

It gives a concrete decision tree (grant despite warnings = arrogance; cautious refusal = still dangerous) and a named test, but the verbs remain abstract ('observe', 'monitor', 'interpret') without concrete operational steps, fitting 'some concrete guidance but incomplete'.

2 / 3

Workflow Clarity

Steps are clearly sequenced (1-4) with an explicit interpretation branch, and the Verification section provides explicit validation checkpoints, matching the anchor for a clear sequence with validation.

3 / 3

Progressive Disclosure

At under 50 lines with a single task and no need for external references, the well-organized sections (Overview, Steps, Example, Expected Outcomes, Verification) satisfy the simple-skill allowance for a top score.

3 / 3

Total

11

/

12

Passed

Description

85%

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is strong: it states concrete capabilities, an explicit trigger, and a distinctive niche. Its only weakness is trigger-term quality, where the language is somewhat specialized rather than mirroring the natural phrasings a user would utter.

Suggestions

Add more colloquial trigger phrasings (e.g., 'assess whether an adversary is overconfident', 'test if enemy leaders ignore good advice') alongside the formal language.

Include a couple of common keyword variations users might say, such as 'enemy overconfidence' or 'reckless enemy leadership'.

DimensionReasoningScore

Specificity

Lists several concrete actions—"make a resource request", "observe whether wise counsel is heeded or ignored", "interpret the response to assess arrogance and internal divisions for attack timing"—matching the anchor for multiple specific concrete actions.

3 / 3

Completeness

An explicit "Use when..." clause states the trigger, and the rest clearly states what the skill does (run a resource-request test, observe counsel heeding, interpret the response), satisfying both what and when.

3 / 3

Trigger Term Quality

It has relevant phrases like "enemy leadership", "judgment quality", "decision-making", and "arrogance", but the vocabulary is fairly niche and lacks common variations a user would naturally say, fitting the 'some relevant keywords but missing common variations' anchor.

2 / 3

Distinctiveness Conflict Risk

It carves a clear niche—strategic reconnaissance via a specific request-based test to gauge enemy arrogance and attack timing—unlikely to conflict with or trigger for other skills.

3 / 3

Total

11

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
baojie/shiji-kb
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.