CtrlK
BlogDocsLog inGet started
Tessl Logo

testing-mwaa-workflow

Tests Amazon MWAA workflow execution end-to-end: trigger a run and monitor it to completion for Provisioned (Python DAG, via Airflow REST API) and Serverless (YAML workflow, via StartWorkflowRun). Verifies the artifact is deployed and parse-ready, triggers with confirmation, polls to terminal state, and on failure delegates diagnosis to debugging-mwaa-workflow and artifact/redeploy fixes to authoring-mwaa-workflow, then retests up to a capped number of attempts. Triggers on: test my DAG, test my workflow, run my DAG, trigger a test run, does my DAG work, smoke-test the pipeline, verify my workflow runs, execute my DAG to check it. Not applicable to writing or deploying a new workflow (handled by authoring-mwaa-workflow), or for diagnosing why a run failed or root-causing an error (handled by debugging-mwaa-workflow).

78

Quality

98%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary process skill body: an unambiguous step spine, explicit validation and safety gates, a well-defined failure-delegation and retest loop, and correct offloading of exact commands to real one-level-deep references. The only deductions are minor: some retest-hygiene and AF2/AF3 readiness detail is duplicated between the steps and the Troubleshooting table, making the main file slightly longer than a pure overview.

Suggestions

Deduplicate the retest-hygiene rules stated in Step 2 and again in Step 5b, and the AF2/AF3 is_active/is_stale readiness logic stated in Step 1 and again in the Troubleshooting table — state each once and cross-reference it.

Consider moving the detailed AF2/AF3 field semantics (last_parsed_time polling window, RestApiClientException-on-POST diagnosis) into references/provisioned-testing.md, keeping Step 1 to the decision rule, to shorten the main file.

Tighten the Step 3 timeout prose by collapsing the three flavor cases into the existing table format alongside the poll-interval table.

DimensionReasoningScore

Conciseness

The body is dense operational guidance with no concept explanations Claude already knows — it stays on non-obvious specifics like "On AF3 (REST API v2) the is_active field does not exist — require is_stale: false instead". Minor trimmable redundancy keeps it below the lean anchor: retest hygiene is stated in Step 2 ("On a retest, all run identifiers must be fresh...") and restated in 5b, and the Troubleshooting rows re-explain Step 1 AF2/AF3 readiness logic already covered inline. Not 5 because of that duplication; not 3 because there is no padded or over-explanatory material.

4 / 5

Actionability

Concrete, executable specifics throughout: exact CLI calls ("aws mwaa get-environment", "get-workflow --workflow-arn"), endpoints ("the /dags collection endpoint", "PATCH /dags/<dag_id> {\"is_paused\": false}"), field names ("last_parsed_time", "has_import_errors", "is_stale"), numeric thresholds ("default to 300s", "15s intervals", "default 3"), and a copy-paste report template, with the exact command sequences correctly carried by the two real reference files. It covers the common Provisioned and Serverless cases the way the fully-executable anchor requires; it is not 4 because there is no pseudocode or missing key detail.

5 / 5

Workflow Clarity

Steps 0–6 are explicitly sequenced with validation checkpoints at every risky point: a mandatory readiness check before triggering ("Do not trigger blind"), a confirmation safety gate with prod re-approval at every state-changing step, a freshness gate before adopting prior runs, terminal-state confirmation before re-triggering, and an explicit debug → fix → retest feedback loop with a progress-based cap and enumerated stop conditions. This is the clear-sequence-with-explicit-validation-and-recovery-loops anchor; it exceeds 4 because validation is not merely present but explicitly mandatory and repeated at each loop iteration.

5 / 5

Progressive Disclosure

Decision logic lives inline while exact commands are properly split into two real, well-signaled, one-level-deep references (references/provisioned-testing.md and references/serverless-testing.md, both present in the bundle, each summarized in the References section and linked contextually in Step 1). Not 5 because the SKILL.md body itself runs ~280 lines and carries some content that could equally live in the references (the full AF2/AF3 readiness field semantics and the RestApiClientException rows duplicated between Step 1 and the Troubleshooting table), leaving minor organization gaps; not 3 because the split is clean, clearly signaled, and easy to navigate.

4 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A model description: third-person voice, comprehensive concrete actions, an explicit natural-language trigger list, and explicit non-applicability boundaries that resolve conflicts with sibling authoring/debugging skills. It is long but every clause is functional with no fluff or over-claims.

DimensionReasoningScore

Specificity

"Tests Amazon MWAA workflow execution end-to-end: trigger a run and monitor it to completion for Provisioned (Python DAG, via Airflow REST API) and Serverless (YAML workflow, via StartWorkflowRun)" plus "Verifies the artifact is deployed and parse-ready... polls to terminal state... delegates diagnosis... retests up to a capped number of attempts" lists multiple specific concrete actions covering the full workflow with no filler. Every clause names a concrete operation, matching the comprehensive-coverage anchor rather than the 'several specific actions, minor gaps' anchor below it.

5 / 5

Completeness

It explicitly answers what ("trigger a run, poll it to a terminal state... capped debug -> fix -> retest loop" behaviors enumerated), when (a literal "Triggers on:" list of concrete phrases), and even when-not ("Not applicable to writing or deploying a new workflow... or for diagnosing why a run failed"), which exceeds the clear what+when anchor. Not 4 because the 'when' is fully explicit with concrete trigger phrases, not merely present-but-imprecise.

5 / 5

Trigger Term Quality

"Triggers on: test my DAG, test my workflow, run my DAG, trigger a test run, does my DAG work, smoke-test the pipeline, verify my workflow runs, execute my DAG to check it" provides comprehensive natural phrasings users would actually say, including question forms ("does my DAG work"), imperative synonyms ("run"/"trigger"/"execute"/"test"), and the colloquial "smoke-test the pipeline" — matching the synonym-inclusive top anchor, not the 'a few natural terms missing' anchor at 4.

5 / 5

Distinctiveness Conflict Risk

The MWAA-testing niche is unambiguous, and the description actively disambiguates from sibling skills ("delegates... diagnosis to debugging-mwaa-workflow and artifact/redeploy fixes to authoring-mwaa-workflow"; "Not applicable to writing or deploying a new workflow (handled by authoring-mwaa-workflow)"), giving minimal conflict risk. This matches the 'clear niche with distinct triggers' anchor rather than the 'minor overlap risk with closely related skills' anchor at 4, since the overlaps are explicitly resolved by name.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
aws/agent-toolkit-for-aws
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.