CtrlK
BlogDocsLog inGet started
Tessl Logo

testing-enemy-arrogance

Use when probing enemy leadership's judgment quality and decision-making state. Based on Zhong's grain-request test against Wu: make a resource request, observe whether wise counsel is heeded or ignored, and interpret the response to assess arrogance and internal divisions for attack timing.

57

Quality

65%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

Fix and improve this skill with Tessl

tessl review fix ./kg/ontology/ontology-v1/skus/procedural/skill_109/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is well-structured, concise, and clearly sequenced for a simple instruction-only skill. Its main weakness is actionability: the steps describe what to do in the abstract without concrete, executable guidance on how to run the test.

Suggestions

Make steps executable: specify the concrete form of the resource request, what to observe in the decision process, and how to record the advisor-counsel signal rather than 'Monitor decision process'.

Add 1-2 worked example scenarios (beyond the Zhong/Wu citation) showing the request made and the interpretation logic applied, so Claude can pattern-match the procedure.

DimensionReasoningScore

Conciseness

The body is brief and largely free of padding, assuming Claude's competence and mostly letting steps speak for themselves; only minor wording could be trimmed.

4 / 5

Actionability

Steps are high-level hints ('Make a request that tests generosity', 'Monitor decision process') with no concrete executable procedure, commands, or specific examples of how to actually conduct the test.

2 / 5

Workflow Clarity

The four steps are clearly sequenced and the Interpret-results step plus a dedicated Verification section provide explicit checkpoints, though the operation is analytical rather than destructive and validation is interpretive rather than automated.

4 / 5

Progressive Disclosure

This is a short single-purpose skill under 50 lines with no bundle files; it is well-organized into Overview, Steps, Example, Expected Outcomes, and Verification sections, fitting the simple-skill exception for a top score.

5 / 5

Total

15

/

20

Passed

Description

67%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly answers both what the skill does and when to use it, and frames a concrete test procedure. It loses points mainly because the trigger terms are domain-narrative rather than natural user phrases, and the scope is somewhat narrow.

Suggestions

Add natural user-facing trigger phrases (e.g. 'Use when planning an attack and needing to gauge enemy overconfidence, arrogance, or whether leadership ignores advisors') to broaden real-world keyword coverage.

Include synonym triggers like 'enemy arrogance test', 'leadership judgment probe', or 'advisor-counsel test' so the skill matches varied phrasings.

DimensionReasoningScore

Specificity

Names the domain ('probing enemy leadership's judgment quality') and several concrete actions ('make a resource request, observe whether wise counsel is heeded or ignored, and interpret the response'), with only minor gaps in coverage.

4 / 5

Completeness

It states both 'what' (probing leadership judgment via a grain-request test) and 'when' ('Use when probing enemy leadership's judgment quality and decision-making state'), with the 'when' explicit but somewhat narrow in scope.

4 / 5

Trigger Term Quality

It includes a natural 'Use when' clause, but the trigger phrases are domain-specific narrative terms ('enemy leadership's judgment quality', 'attack timing') rather than the everyday keywords a user would naturally utter, missing common synonyms.

3 / 5

Distinctiveness Conflict Risk

The niche is fairly distinct (a specific historical-test reconnaissance method tied to attack timing), with only minor overlap risk against other generic military-strategy skills.

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
baojie/shiji-kb
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.