CtrlK
BlogDocsLog inGet started
Tessl Logo

agents-optimize

Use when measuring or improving agent quality and performance — set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization. Triggers on: "evaluate my agent", "add evaluator", "measure quality", "quality gate", "run evals", "agent too slow", "why is it slow", "reduce latency", "set up observability", "CloudWatch dashboard", "how much does my agent cost", "cost optimization", "logs not showing up", "logs missing", "spans not found", "eval failing", "eval error", "dev traces", "local traces", "agentcore dev traces", "traces to CloudWatch". Not for debugging errors or crashes — use agents-debug. Slow but correct routes here; broken routes to debug.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured router skill: concise overview, deterministic workflow routing to real one-level-deep reference files, and early gating checks. The body delegates all executable procedure to the references, which keeps it token-efficient but leaves it just short of fully self-contained actionability and in-body validation.

Suggestions

Trim or merge the 'When to use' bullets with the routing table in Step 2 — they duplicate the frontmatter description's triggers and could be cut without losing information.

Replace the filler Step 3 ('The reference file contains the full procedure. Follow it step by step.') and the 'Output: Depends on the workflow' section with a one-line note embedded in Step 2, reclaiming tokens.

Add an explicit in-body validation checkpoint after the reference workflow completes (e.g., 'verify the eval ran and produced a score before declaring success') so the top-level sequence contains a feedback loop rather than delegating all verification to the references.

DimensionReasoningScore

Conciseness

The body is lean and assumes competence (e.g., a compact routing table instead of prose), but there is minor redundancy: the 'When to use' bullets restate the frontmatter triggers, and Step 3 ("The reference file contains the full procedure. Follow it step by step.") plus "Output: Depends on the workflow" are near-contentless filler. This matches 'efficient; minor instances of over-explanation that could be trimmed' rather than the every-token-earns-its-place 5 anchor.

4 / 5

Actionability

Concrete, executable guidance dominates: a specific version command (`agentcore --version`), a specific context file to read (`agentcore/agentcore.json`) with a ready-to-use failure message, and an intent-to-reference routing table with markdown links. It stops short of fully copy-paste-ready procedure in the body itself — the actual commands live in the references — so it fits 'mostly executable guidance with minor gaps' rather than the 5 anchor.

4 / 5

Workflow Clarity

Steps 0-3 are clearly sequenced with two gating checkpoints (CLI version floor, project-file existence check with an explicit halt message), and the routing table makes branch selection deterministic. However, no validation of outcomes occurs in the body (success verification is delegated to the references), matching 'clear sequence with most checkpoints present; minor validation gaps' rather than the 5 anchor with explicit validation steps and feedback loops.

4 / 5

Progressive Disclosure

The body is a clean overview router: three one-level-deep references (evals.md, observability.md, cost.md — all verified present on disk), each clearly signaled with markdown links in an intent-mapped table, with detailed procedure appropriately split out and easy navigation. This matches the top anchor ('clear overview with well-signaled one-level-deep references; content appropriately split').

5 / 5

Total

17

/

20

Passed

Description

91%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: explicit what-and-when structure, an unusually rich set of natural trigger phrases with synonyms, and explicit boundary routing to sibling skills. The only weaknesses are that the stated actions are capability domains rather than fully concrete operations, and a handful of generic observability/logging triggers that slightly raise conflict risk.

DimensionReasoningScore

Specificity

The description enumerates five concrete capability areas ("set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization") which is broad coverage, but these are capability domains rather than fully concrete operations (e.g., no specific commands or artifacts), so it sits between the 'several specific actions' and 'comprehensive concrete actions' anchors — noticeably above the midpoint, hence 4.

4 / 5

Completeness

Both halves are explicit: 'what' is stated up front ("measuring or improving agent quality and performance — set up evaluators, online monitoring, CI/CD quality gates, observability, or cost optimization") and 'when' is given via a leading "Use when" clause plus 21 concrete trigger phrases. This matches the top anchor ('clearly and explicitly answers both what AND when with concrete trigger phrases'); the 4 anchor's 'when could be more explicit' is not the case here.

5 / 5

Trigger Term Quality

The trigger list is comprehensive and includes natural user phrasings with synonyms and variations: "evaluate my agent", "agent too slow", "why is it slow", "how much does my agent cost", "logs not showing up" / "logs missing", "spans not found", "eval failing" / "eval error", "dev traces" / "local traces" / "agentcore dev traces". This matches the anchor for comprehensive coverage including synonyms; the score-4 anchor's 'a few natural terms missing' does not apply.

5 / 5

Distinctiveness Conflict Risk

The description has a clear niche (agent quality/performance/observability) and explicit disambiguation ("Not for debugging errors or crashes — use agents-debug. Slow but correct routes here; broken routes to debug"), but several triggers ("set up observability", "logs not showing up", "reduce latency") are generic outside the agent context and carry minor overlap risk with other ops/logging skills — matching 'mostly distinct; minor overlap risk' rather than the clear-niche 5 anchor.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
aws/agent-toolkit-for-aws
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.