CtrlK
BlogDocsLog inGet started
Tessl Logo

agents-harden

Use when preparing your agent for production — IAM scoping, inbound auth (JWT, SigV4), secrets management, cold start optimization, session lifecycle, rate limiting, input validation, and quota guidance. Triggers on: "production checklist", "harden agent", "production ready", "secure agent", "inbound auth", "going live", "cold start optimization", "session lifecycle", "StopRuntimeSession", "quota", "throttling", "maxVms", "rate limit", "security audit of outbound API calls", "gateway target audit for production", "restrict who can call", "lock down endpoint", "only our app can call". Not for Cedar tool-restriction policies — use agents-connect. Not for quality measurement — use agents-optimize. Not for outbound credential storage or API key wiring — use agents-connect. Not for A2A agent-to-agent auth — use agents-build. Cold start observation and diagnosis (not optimization) routes to agents-debug.

72

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-sequenced production-hardening guide with excellent executable commands, policies, and diagnostic loops, and a properly offloaded limits reference. Its main weakness is token efficiency: at ~680 lines it carries tighten-able prose and an inline deprecated-guidance aside, and two large sections are reference-sized material inlined in SKILL.md.

Suggestions

Move the cold start optimization and session lifecycle sections (~155 lines combined) into reference files like references/cold-start.md and references/sessions.md, linked from a short summary section — limits.md already demonstrates this pattern well.

Delete or relocate the 'The skill previously recommended CodeZip over Container' changelog aside into a deprecated/old-patterns note; it is skill-history rather than guidance for the current reader.

Tighten the prose in 'Where cold start time actually goes' and the session-lifecycle trade-off narrative into the existing tables and bullet lists — much of the surrounding explanation restates what the tables already show.

DimensionReasoningScore

Conciseness

The body is dense and platform-specific rather than teaching generic concepts, but at ~680 lines it includes tighten-able material: the inline changelog aside 'The skill previously recommended CodeZip over Container when possible. That's an oversimplification', and multi-paragraph prose in the cold-start and session-lifecycle sections. It fits 'mostly efficient but includes some unnecessary explanation or could be tightened' — above level 2 (no padding or generic library intros) but below level 4 (more than minor instances of over-explanation).

3 / 5

Actionability

Fully executable throughout: copy-paste bash ('agentcore status --json | jq -r ".runtimes[0].executionRoleArn"', the EventBridge put-rule command), complete JSON IAM/trust policies, runnable Python (credential decorators, async-task registration, deferred client init), and a hit-to-action decision table for outbound HTTP calls. Specific examples cover the common production cases, matching the level-5 anchor.

5 / 5

Workflow Clarity

The process is explicitly sequenced (verify CLI version → read agentcore/agentcore.json → work through each checklist category), with verification commands per category, ordered diagnostic feedback loops (the 5-step JWT 403 walkthrough 'Walk through these in order', the 4-step maxVms remediation ending 'Only then ... request an increase'), and a final production checklist — matching the level-5 anchor of explicit validation, feedback loops, and checklists.

5 / 5

Progressive Disclosure

references/limits.md (291 lines) is properly offloaded, linked three times with a clear summary of what it covers, and sibling-skill pointers are named and one level deep. However, the body inlines reference-sized sections (cold start ~70 lines, session lifecycle ~85 lines) that could be split out, keeping it at 'good structure; minor organization gaps' rather than the level-5 'content appropriately split'.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: it states concrete capabilities, gives comprehensive natural trigger phrases, explicitly covers both what and when, and disambiguates against four sibling skills with explicit routing. No changes needed.

DimensionReasoningScore

Specificity

Lists eight concrete capability areas — 'IAM scoping, inbound auth (JWT, SigV4), secrets management, cold start optimization, session lifecycle, rate limiting, input validation, and quota guidance' — with specific technology names, comprehensively covering the production-hardening domain. It exceeds level 4 (which allows minor coverage gaps) and is far beyond naming 1-2 actions.

5 / 5

Completeness

Explicitly answers both what (the concrete capability list) and when ('Use when preparing your agent for production' plus 'Triggers on:' with concrete phrases). This matches the level-5 anchor exactly; level 4 would require the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

Comprehensive natural trigger phrases including synonyms ('production checklist', 'production ready', 'going live', 'harden agent') and concrete API identifiers ('StopRuntimeSession', 'maxVms', 'throttling', 'rate limit'). Coverage includes the variations users would naturally say, meeting the level-5 anchor rather than the 'a few natural terms missing' level 4.

5 / 5

Distinctiveness Conflict Risk

Four explicit negative boundaries ('Not for Cedar tool-restriction policies — use agents-connect', 'Not for quality measurement — use agents-optimize', 'Not for A2A agent-to-agent auth — use agents-build', 'Cold start observation and diagnosis ... routes to agents-debug') give it a clear niche with minimal conflict risk, exceeding the 'minor overlap risk' of level 4.

5 / 5

Total

20

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (705 lines); consider splitting into references/ and linking

Warning

relative_links

Relative link issues: 5 suspicious

Warning

referenced_paths_exist

Referenced path issues: 1 missing

Warning

Total

13

/

16

Passed

Repository
aws/agent-toolkit-for-aws
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.