CtrlK
BlogDocsLog inGet started
Tessl Logo

aws-observability

Builds, configures, debugs, and optimizes AWS observability — operator-symptom questions and detecting Omni vs classic CloudWatch. CloudWatch: Log Insights, alarms, Dynamic Instrumentation, and Application Signals — instrumenting/onboarding a service to Application Signals with ADOT on EC2/ECS/EKS/Lambda: auto-instrumentation, monitored service, reporting telemetry, ServiceEvents, CI/CD metadata, Terraform/manifest. Also fleet health views. CloudWatch Omni on an existing Space: SQL over logs and traces, PromQL over metrics, Omni dashboards, Omni alerts, context graph for root cause, programmatic/IaC access (API/SDK/CLI/CloudFormation) and driving Omni from a coding agent or skills, and evaluating AI agent quality from traces — on-demand and continuous online scoring of live agent traffic, readback, and custom trace evaluators. For first-time Omni setup — creating a Space, granting access, ingestion, or ADOT instrumentation — use setting-up-cloudwatch-observability. Not for app logging or threat detection.

68

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary routing-oriented SKILL.md: a two-product decision tree with probe commands, error-recovery paths, confirmation gates for destructive operations, and output-contract checklists, all backed by a well-organized one-level-deep reference bundle. The one real weakness is conciseness — the body is long and repeats the must-state instruction pattern many times, and some of the denser exception logic would fit better in a reference file.

Suggestions

Consolidate the repeated 'open X and surface every item in its facts-you-MUST-surface checklist' pattern into a single stated convention (e.g. one rule in the routing intro: 'whenever you route to a file, surface its MUST-state checklist') instead of restating it per section and per routing row.

Move the long inline exception chains (e.g. the alarm/PromQL ambiguity exception and the 2a/2b/2c sub-rule detail) into a short disambiguation reference file, keeping only the trigger terms and the resulting route in SKILL.md.

Trim duplicate guidance between the Step 0 numbered rules and the 'Under-specified alert requests' / 'Under-specified dashboard requests' paragraphs, which restate routing details already covered by the routing tables.

DimensionReasoningScore

Conciseness

Mostly efficient — the body is dense domain-specific routing knowledge Claude does not already know (Omni vs CloudWatch split, Region-per-Space trap, client-version probe failure handling), not padding, and it correctly defers detail to reference files. But it could be tightened: the "surface every item in its 'facts you MUST surface' checklist" instruction is repeated across at least six sections, and several rules carry long exception chains (e.g. the alarm/PromQL exception in rule 1) that could be consolidated into a reference file.

3 / 5

Actionability

Fully actionable for an instruction-only skill: copy-paste probe commands ("aws cloudwatchomni list-domains", "aws cloudwatchomni list-spaces"), concrete upgrade commands ("brew upgrade awscli", "pip install -U boto3 botocore"), per-need routing tables naming the exact file to open, and named scripts ("scripts/cloudwatch-omni/evaluate_traces.py", "scripts/cloudwatch/di_instrumentation.py"). Every routing row resolves to a specific, verified action.

5 / 5

Workflow Clarity

The decision workflow is clearly sequenced (Step 0 product decision → Step 0.5 symptom routing → per-product routing tables) with explicit validation gates and feedback loops: probe errors route to a client-upgrade-and-retry path, Dynamic Instrumentation (a destructive, live-service operation) requires confirmation before any create/delete and a narrate-before-acting loop, dashboards mandate validate-before-save, and "Still inconclusive → ask the customer" provides an explicit terminal fallback. Checklists (must-state output contracts) are attached to every complex output type.

5 / 5

Progressive Disclosure

SKILL.md is a true overview/router: nearly all detail lives in one-level-deep reference files (references/cloudwatch/*, references/cloudwatch-omni/*, query/ subdirectory, scripts/, assets/), every referenced path in the body resolves to a real bundle file, and a "Files" section catalogues each reference with a one-line content description. The only second-level hop (dynamic-instrumentation.md → dynamic-instrumentation/) is explicitly signaled in the Files table, so navigation stays easy.

5 / 5

Total

18

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A dense, highly specific description that comprehensively names the skill's actions and cleanly separates two products and a sibling setup skill. Its main weaknesses are the absence of a positive "Use when..." trigger clause (when-guidance is expressed as scope conditions and exclusions instead) and a few missing natural symptom phrases that operators would use.

DimensionReasoningScore

Specificity

Lists many specific concrete actions across the domain: "Builds, configures, debugs, and optimizes", "Log Insights, alarms, Dynamic Instrumentation", "instrumenting/onboarding a service to Application Signals with ADOT on EC2/ECS/EKS/Lambda", "SQL over logs and traces, PromQL over metrics", "context graph for root cause", "programmatic/IaC access (API/SDK/CLI/CloudFormation)", and "on-demand and continuous online scoring of live agent traffic". This is comprehensive concrete-action coverage, matching the anchor-5 example's breadth; no anchor-4-style gaps in coverage are evident.

5 / 5

Completeness

The "what" is explicit and detailed, and the "when" is present through scope conditions — "CloudWatch Omni on an existing Space", "For first-time Omni setup ... use setting-up-cloudwatch-observability", and "Not for app logging or threat detection" — so it exceeds anchor 3 (when only weakly implied). But there is no positive "Use when..." trigger clause; when-guidance is fragmented into scope/exclusion statements and could be more explicit, which is exactly the anchor-4 description.

4 / 5

Trigger Term Quality

Good keyword coverage with natural terms users would say: "observability", "logs", "traces", "metrics", "alarms", "alerts", "dashboards", "query", "root cause", "CI/CD metadata", "Terraform". However, common symptom phrasings operators actually use ("high latency", "error rate", "throttled", "is my service healthy", "monitoring") are absent even though the description gestures at them via "operator-symptom questions", so a few natural terms are missing — anchor 4 rather than 5.

4 / 5

Distinctiveness Conflict Risk

A clear niche with distinct triggers: AWS observability split across two explicitly named products ("Omni vs classic CloudWatch"), with explicit disambiguation from the sibling skill ("use setting-up-cloudwatch-observability" for setup) and negative scoping ("Not for app logging or threat detection"). Terms like "Space", "Log Insights", and "gen_ai"-style evaluation vocabulary are unlikely to collide with other skills — minimal conflict risk.

5 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 69 deeper-than-1-level

Warning

referenced_paths_exist

Referenced path issues: 74 deeper-than-1-level

Warning

Total

14

/

16

Passed

Repository
aws/agent-toolkit-for-aws
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.