CtrlK
BlogDocsLog inGet started
Tessl Logo

agents-debug

Use when your agent or environment is broken — wrong answers, errors, timeouts, tool failures, or CLI issues. Reads traces and logs to diagnose root causes. Also checks prerequisites when the CLI itself isn't working. Triggers on: "agent not working", "wrong answer", "agent error", "tool call failing", "debug agent", "check logs", "read traces", "broken", "500 error", "424 error", "model access denied", "command not found", "stuck in DELETING", "maxVms exceeded", "cold start diagnosis", "cold start slow", "agentcore create error", "create failed", "exit code 7", "connection refused local dev". Not for deploy failures — use agents-deploy. Not for performance tuning without errors — use agents-optimize. Not for VPC configuration — use agents-build. Not for observability setup or missing logs — use agents-optimize.

65

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

This is a highly actionable, well-structured diagnostic guide: concrete commands, copy-paste code, and a clear triage-to-symptom workflow with per-symptom verification steps. Its weaknesses are efficiency (duplicated latency guidance, overlong code examples) and progressive disclosure — a very large body that inlines several topics that should live in reference files like the existing doctor.md.

Suggestions

Deduplicate the CloudWatch ingestion-latency guidance: state the ~10s/~15s wait once (either Step 3 or the 'No traces appearing' section) and cross-reference it from the other, and drop the meta-commentary about what older docs said.

Move self-contained deep-dive topics into reference files (e.g., references/streaming.md for the keepalive pattern and client filtering, references/iam-logging.md for the IaC IAM policy JSON) and keep one-line pointers in SKILL.md, mirroring how doctor.md is already handled.

Trim the emit_keepalive example to the minimal pattern (~10 lines) and fix the repeated "1." list numbering in the 'Memory not working' section so the rendered sequence is correct.

DimensionReasoningScore

Conciseness

Mostly efficient and dense with platform-specific facts Claude wouldn't know (maxVms quota semantics, port auto-increment, OTEL wrapping), but there is tightening to do: the ~10s CloudWatch latency guidance appears twice (the "Important" paragraph in Step 3 and the "Symptom: No traces appearing" section), the Step 3 paragraph spends tokens on stale-docs meta-commentary ("Older skills and docs said 30–60s ... both are stale"), and the ~30-line keepalive example plus client-side filter snippet run long. This fits 'mostly efficient but includes some unnecessary explanation or could be tightened' rather than the minor-trim 4 anchor.

3 / 5

Actionability

Fully executable throughout: exact CLI commands ("agentcore traces list --runtime <AgentName> --since 1h", "aws logs tail /aws/lambda/<function-name>"), copy-paste code (stop_runtime_session call, emit_keepalive streaming pattern, IAM policy JSON), and concrete ordered fixes for each symptom. Specific examples cover the common cases, matching the top anchor.

5 / 5

Workflow Clarity

The intake workflow is clearly sequenced (Step 0 problem-type triage → verify CLI version → classify symptom → read traces/logs → symptom-specific diagnosis), and each symptom section follows a check-then-fix order with verification commands (e.g., "If still no traces after ~30 seconds: 1. Verify observability... 2. Check the agent was actually invoked... 3. Check CloudWatch permissions"). It falls short of the 5 anchor because a few fixes lack an explicit post-fix verification step (e.g., the port-kill fix `lsof -tiTCP:8080 ... | xargs kill` and the IAM redeploy) and the Memory section has broken list numbering (repeated "1." items).

4 / 5

Progressive Disclosure

Section structure is good and the one bundle reference ([references/doctor.md]) is real, well-signaled, and one level deep, but the body is a ~700-line monolith while the bundle contains only that single file — deep-dive material that clearly belongs in separate reference files is inlined (the streaming keepalive pattern with full Python code, the IAM policy JSON for IaC deploys, the cross-region inference profile table, and the LangGraph/X-Ray specifics). This matches 'some structure but could be better organized; content that should be separate is inline' rather than the 4 anchor, where most content would be appropriately placed across the bundle.

3 / 5

Total

15

/

20

Passed

Description

90%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it explicitly states what the skill does and when to use it, carries an unusually rich set of natural trigger phrases (including niche error signatures), and explicitly routes adjacent use cases to sibling skills to avoid mis-triggering. The only soft spot is that the capability list itself is thin relative to the trigger list.

DimensionReasoningScore

Specificity

Names the domain (agent/environment debugging) and a couple of concrete actions — "Reads traces and logs to diagnose root causes" and "checks prerequisites when the CLI itself isn't working" — but does not enumerate several specific actions, leaning instead on a long trigger list. This matches the 'names domain and 1-2 concrete actions, but not comprehensive' anchor; a 4 would require a fuller list of distinct capabilities rather than trigger phrases.

3 / 5

Completeness

Explicitly answers both: what ("Reads traces and logs to diagnose root causes. Also checks prerequisites") and when ("Use when your agent or environment is broken — wrong answers, errors, timeouts, tool failures, or CLI issues" plus a concrete trigger list). Both halves are stated clearly with concrete trigger phrases, matching the 5 anchor exactly.

5 / 5

Trigger Term Quality

"Triggers on" lists twenty natural phrases including synonyms and specific error signatures users would actually say: "agent not working", "wrong answer", "check logs", "command not found", "stuck in DELETING", "exit code 7", "connection refused local dev". This comprehensively covers natural terms and variants, matching the top anchor; nothing significant is missing for the domain.

5 / 5

Distinctiveness Conflict Risk

It carves out a clear niche with distinct triggers and defuses overlap via explicit negative boundaries — "Not for deploy failures — use agents-deploy. Not for performance tuning without errors — use agents-optimize. Not for VPC configuration — use agents-build. Not for observability setup or missing logs — use agents-optimize." Minimal conflict risk, matching the top anchor.

5 / 5

Total

18

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (722 lines); consider splitting into references/ and linking

Warning

relative_links

Relative link issues: 11 suspicious

Warning

referenced_paths_exist

Referenced path issues: 11 missing

Warning

Total

13

/

16

Passed

Repository
aws/agent-toolkit-for-aws
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.