CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-orchestration-improve-agent

Systematic improvement of existing agents through performance analysis, prompt engineering, and continuous iteration.

24

Quality

13%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./plugins/antigravity-awesome-skills/skills/agent-orchestration-improve-agent/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

12%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This skill is a verbose, high-level process document that reads more like a textbook chapter on agent optimization than an actionable skill for Claude. It explains many concepts Claude already knows, uses placeholder pseudocode instead of executable examples, and packs everything into a single long file. While the overall workflow structure is reasonable, the lack of concrete, executable guidance and the excessive verbosity significantly reduce its utility.

Suggestions

Replace pseudocode blocks with real, executable commands or code snippets that Claude can actually run — define what 'context-manager' and 'prompt-engineer' tools are and how to invoke them, or remove those references entirely.

Cut the content by at least 60% by removing explanations of concepts Claude already knows (A/B testing, semantic versioning, Cohen's d, chain-of-thought prompting) and keeping only project-specific decisions and configurations.

Split detailed sections (evaluation metrics, testing framework, deployment procedures) into separate referenced files to improve progressive disclosure and reduce the main file to a concise overview.

Add concrete, real-world examples with actual inputs and outputs rather than template placeholders with '[X%]' brackets.

DimensionReasoningScore

Conciseness

Extremely verbose at ~300+ lines. Explains many concepts Claude already knows (what chain-of-thought is, what A/B testing is, what semantic versioning means, what Cohen's d is). Extensive bullet-point lists describe general best practices rather than providing novel, specific guidance. Much of this content is generic software engineering and prompt engineering knowledge that doesn't earn its token cost.

1 / 3

Actionability

Despite its length, the skill contains no executable code or concrete commands. The code blocks are pseudocode placeholders (e.g., 'Use: context-manager', 'Use: prompt-engineer') referencing tools that aren't defined or documented. Metrics templates use placeholder brackets like '[X%]'. There are no copy-paste-ready examples, real tool invocations, or specific implementation details.

1 / 3

Workflow Clarity

The four-phase structure provides a clear sequence (analyze → improve → test → deploy), and rollback triggers with thresholds are specified. However, validation checkpoints between phases are implicit rather than explicit, and the feedback loop between testing failure and re-optimization is not clearly articulated as a concrete decision gate. The rollback section is good but the overall workflow lacks explicit 'stop and verify before proceeding' gates between phases.

2 / 3

Progressive Disclosure

All content is inlined in a single monolithic file with no references to supporting documents. The skill is extremely long and would benefit enormously from splitting detailed sections (evaluation metrics, testing frameworks, deployment procedures) into separate referenced files. No bundle files are provided, and no external references are made.

1 / 3

Total

5

/

12

Passed

Description

14%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

This description relies heavily on abstract buzzwords ('systematic improvement', 'continuous iteration') without specifying concrete actions or when the skill should be triggered. It lacks a 'Use when...' clause and is too generic to be reliably distinguished from other skills related to prompt engineering, agent development, or code optimization.

Suggestions

Add a 'Use when...' clause with specific trigger scenarios, e.g., 'Use when the user wants to improve an existing agent's performance, debug agent behavior, optimize prompts, or run evaluations.'

Replace vague phrases with concrete actions, e.g., 'Analyzes agent outputs against expected results, rewrites system prompts, adds few-shot examples, and designs evaluation benchmarks.'

Include natural trigger terms users would say, such as 'optimize agent', 'improve prompts', 'agent not working', 'eval results', 'prompt tuning', or 'agent debugging'.

DimensionReasoningScore

Specificity

The description uses vague, abstract language like 'systematic improvement', 'performance analysis', and 'continuous iteration' without listing concrete actions. These are buzzwords rather than specific capabilities.

1 / 3

Completeness

The 'what' is vaguely stated and the 'when' is entirely missing. There is no 'Use when...' clause or equivalent explicit trigger guidance, which per the rubric should cap completeness at 2, but since the 'what' is also weak, this scores a 1.

1 / 3

Trigger Term Quality

Contains some relevant keywords like 'agents', 'prompt engineering', and 'performance analysis' that users might mention, but misses common variations like 'optimize prompts', 'debug agent', 'improve accuracy', 'eval', 'benchmarks', or 'agent tuning'.

2 / 3

Distinctiveness Conflict Risk

The description is very generic and could overlap with many skills related to code optimization, debugging, testing, prompt writing, or general agent development. 'Systematic improvement' and 'continuous iteration' are too broad to carve out a clear niche.

1 / 3

Total

5

/

12

Passed

Validation

90%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 10 / 11 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

10

/

11

Passed

Repository
popey/claude-code-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.