CtrlK
BlogDocsLog inGet started
Tessl Logo

upgrading-mwaa-environments

Upgrades an MWAA environment to a newer Airflow version — within 2.x, within 3.x, or across the 2.x-to-3.x boundary. Computes the version-jump path, inserting the 2.11.x stepping-stone and Python-transition step when needed. Chooses an approach by whether run history and the same environment (URL/ARN) must be kept: a new-environment upgrade (blue-green), a rehearsed in-place upgrade validated on a test copy, or a direct in-place upgrade. Runs Ruff scanning and deprecation-warning log scans for 3.x moves, plus Docker validation, batched deployment, and switchover. Saves a resumable upgrade plan for multi-session work. Triggers on: upgrade MWAA, upgrade Airflow, migrate to Airflow 3, MWAA Airflow 3, Airflow 2 to 3, preserve Airflow history, keep same MWAA environment, blue-green cutover, validation environment before cutover. Not for authoring new DAGs (authoring-mwaa-workflow), debugging unrelated DAG failures (debugging-mwaa-workflow), or MWAA Serverless YAML workflows (provisioned Python DAGs only).

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured orchestrator skill: excellent phase sequencing, validation checkpoints, and rollback logic, with a clean one-level-deep reference architecture. Its main weakness is redundancy — the same checkpoint/reload rules are restated across Phases 0, 4, and 7, inflating token cost without adding information.

Suggestions

State the test-run checkpoint rule once (e.g., in plan-materialization.md or Phase 7) and reference it from Phases 0 and 4 with a one-line pointer instead of re-explaining it in full three times.

Merge the Phase 0 reload rule and Plan Reconciliation paragraphs, which repeat the 'never act on a remembered summary' and stale-PASSED consequences; a single consolidated statement would cut several hundred tokens.

Give the concrete pull command for Phase 6 step 2 (e.g., 'docker pull <repo>/<image>:<tag>' with the repo URL) so the image-resolution step is executable rather than directional.

DimensionReasoningScore

Conciseness

Mostly non-obvious, domain-specific content with no concept-explanation padding (it assumes Claude knows Airflow/MWAA), but the test-run checkpoint rule is explained in full detail three times (Phase 0, Phase 4 'Write out every checkpoint', Phase 7) and Phase 0's reload/reconciliation block restates its own rules back-to-back. The redundancy is more than the 'minor instances' of the score-4 anchor.

3 / 5

Actionability

Concrete copy-paste commands ('ruff check --preview --select AIR .', 'aws mwaa invoke-rest-api --name <current-env> --method PATCH --path /dags/<dag_id> --body ...') with specific operational thresholds (64-char --path limit, 10-second invoke-rest-api timeout, 0.2s throttle) and fallbacks. Minor gaps remain — e.g. 'Pull the target-version image (resolve the exact tag at runtime)' gives no command, and many steps delegate execution detail to the reference files.

4 / 5

Workflow Clarity

Nine clearly sequenced phases with explicit validation checkpoints and feedback loops: per-jump confirmation gates, the test-run checkpoint blocking switchover/live-upgrade, drain-then-monitor-one-full-cycle per batch, rollback per batch on failure, Docker validate-then-return-to-Phase-5 loop, and plan reconciliation with ground-truth re-establishment on resume. Destructive operations all require per-step approval.

5 / 5

Progressive Disclosure

SKILL.md stays an orchestrating overview; all nine referenced files exist, are one level deep, are listed with one-line descriptions in 'Reference Documentation', and are linked at their point of use (Phase 1 → discovery-preflight.md, Phase 7 → per-approach strategy files). Deep detail (version matrix, checklists, plan-format contract) is correctly split out.

5 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: concrete third-person actions, comprehensive natural trigger phrases, an explicit when-clause, and clear negative boundaries against sibling skills. Despite its length, every clause is load-bearing rather than padded.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Computes the version-jump path, inserting the 2.11.x stepping-stone and Python-transition step', 'Runs Ruff scanning and deprecation-warning log scans', 'Docker validation, batched deployment, and switchover', 'Saves a resumable upgrade plan' — with comprehensive coverage and no generic filler. Third-person voice throughout, matching the anchor for comprehensive specific concrete actions.

5 / 5

Completeness

Explicitly answers both 'what' (upgrade actions, approach selection, validation, switchover, resumable plan) and 'when' via the dedicated 'Triggers on:' clause, plus scope exclusions. Both are concrete and explicit, matching the top anchor rather than the score-4 'when could be more specific' case.

5 / 5

Trigger Term Quality

The explicit trigger list — 'upgrade MWAA, upgrade Airflow, migrate to Airflow 3, MWAA Airflow 3, Airflow 2 to 3, preserve Airflow history, keep same MWAA environment, blue-green cutover, validation environment before cutover' — covers the natural phrases a user would say, including synonyms, product names, and version-boundary variants.

5 / 5

Distinctiveness Conflict Risk

Narrow MWAA-upgrade niche with distinct triggers, reinforced by an explicit negative boundary — 'Not for authoring new DAGs (authoring-mwaa-workflow), debugging unrelated DAG failures (debugging-mwaa-workflow), or MWAA Serverless YAML workflows' — minimizing conflict risk with sibling skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
aws/agent-toolkit-for-aws
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.