CtrlK
BlogDocsLog inGet started
Tessl Logo

ingesting-into-data-lake

Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB, or existing Glue catalog tables (migration). Default target is S3 Tables; standard Iceberg on a general purpose bucket is supported where S3 Tables is not adopted. Handles one-time loads, recurring pipelines, migrations. Triggers on: import data, load data, ingest, sync database, migrate table, move data to AWS, set up pipeline, ETL, pull from Snowflake, query BigQuery into S3, export DynamoDB, CTAS, convert to Iceberg. Do NOT use for setting up or troubleshooting Glue connections (use connecting-to-data-source), creating empty tables (use creating-data-lake-table), running queries (use querying-data-lake), finding tables by fuzzy name (use finding-data-lake-assets), catalog audit (use exploring-data-catalog), or SaaS platforms like Salesforce, ServiceNow, SAP, MongoDB, Kafka.

75

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong orchestration-style SKILL.md: a clearly sequenced seven-step workflow with mandatory validation and error-recovery paths, lean token-efficient prose, and exemplary progressive disclosure across 25 verified one-level-deep reference files. The only weaknesses are minor — some duplication between the Step 2 routing table and the References section, and body-level steps that are precise directives rather than self-contained executable recipes.

Suggestions

Collapse the duplicate source-reference listing: the Step 2 routing table already links all seven source-specific references, so the References > Source-specific subsection could be dropped or reduced to a single line pointing back to Step 2.

Make Step 6 validation self-executable by inlining the one or two core commands (e.g. an Athena COUNT(*) source-vs-target query) instead of only naming the three checks and delegating entirely to data-quality-validation.md.

DimensionReasoningScore

Conciseness

The body is lean — tables, terse imperatives, and specific gotchas ("All spark.sql.catalog.* config MUST go in --conf job arguments, never in spark.conf.set()") with no explanation of concepts Claude already knows. It falls just short of the score-5 'every token earns its place' anchor because of minor redundancy: the seven source-specific references are listed both in the Step 2 routing table and again in the References section, and a few items (e.g. 'one reference per source' bullet) restate what the table already shows.

4 / 5

Actionability

Concrete executable commands appear where they belong in the orchestration flow ("aws sts get-caller-identity", "aws glue get-connection --name <CONNECTION_NAME> --region <REGION>"), argument routing is specific, and gotchas name exact flags and error strings. It sits at score 4 rather than 5 because most executable substance (job templates, validation commands) is delegated to references while the body's own steps, though precise, are directives rather than copy-paste-ready recipes — acceptable for an orchestration skill but not fully self-contained.

4 / 5

Workflow Clarity

A seven-step numbered workflow with explicit checkpoints (verify credentials in Step 1, confirm connection exists in Step 3, clarify target before writing in Step 4) and a mandatory validation checklist for this batch operation — "Run all three, do not skip: 1. Row count matches... 2. Null check... 3. Spot-check 3-5 sample rows". Error recovery is covered by a troubleshooting table and explicit delegation rules ("Connection failures during ingest delegate back to connecting-to-data-source"), matching the score-5 anchor of clear sequence, explicit validation, and feedback loops for recovery.

5 / 5

Progressive Disclosure

The body is a well-organized overview that keeps execution detail in 25 one-level-deep reference files, all verified to exist, each linked inline where relevant (routing table in Step 2, gotchas, Step 6/7 pointers) and again in a categorized References section (source-specific, cross-cutting, migration-specific, JDBC-specific). No nested references, no inlined bulk content — this matches the score-5 anchor of a clear overview with well-signaled one-level-deep references and easy navigation.

5 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: it states concrete third-person capabilities with a comprehensive source inventory, gives an explicit 'Triggers on' list with natural synonyms, and closes with a 'Do NOT use' clause that routes neighboring intents to named sibling skills. Both the 'what' and 'when' questions are answered explicitly, and conflict risk is actively minimized.

DimensionReasoningScore

Specificity

The description enumerates concrete actions across a comprehensive source list — "Import data into the AWS data lake from S3 files, local uploads, JDBC databases (Oracle, SQL Server, PostgreSQL, MySQL, RDS, Aurora), Amazon Redshift, Snowflake, BigQuery, DynamoDB" — plus concrete target behavior ("Default target is S3 Tables; standard Iceberg on a general purpose bucket is supported") and load patterns ("one-time loads, recurring pipelines, migrations"). This matches the score-5 anchor of multiple specific concrete actions with comprehensive coverage; it is well above score 4 since there are no meaningful coverage gaps, and all verbs are concrete third-person actions rather than the generic naming of score 2-3.

5 / 5

Completeness

Both questions are answered explicitly: 'what' via the enumerated sources, targets, and load patterns, and 'when' via the literal "Triggers on:" clause with concrete trigger phrases, plus an explicit negative-scope "Do NOT use for..." clause. This is the score-5 anchor verbatim in structure (clear what AND when with concrete trigger phrases); a score of 4 would require the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

The explicit trigger list covers natural user phrasings and their synonyms: "import data, load data, ingest, sync database, migrate table, move data to AWS, set up pipeline, ETL, pull from Snowflake, query BigQuery into S3, export DynamoDB, CTAS, convert to Iceberg". This matches the score-5 anchor of comprehensive natural-term coverage including synonyms (load/import/ingest/sync; migrate/convert), and clearly exceeds score 4 ('a few natural terms missing').

5 / 5

Distinctiveness Conflict Risk

The description draws a clear niche (data ingestion into the AWS data lake) and actively defuses conflicts by routing adjacent intents to named sibling skills: "Do NOT use for setting up or troubleshooting Glue connections (use connecting-to-data-source), creating empty tables (use creating-data-lake-table), running queries (use querying-data-lake), finding tables by fuzzy name (use finding-data-lake-assets), catalog audit (use exploring-data-catalog)". Minimal conflict risk, matching the score-5 anchor; score 4's 'minor overlap risk' does not apply given the explicit disambiguation.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
aws/agent-toolkit-for-aws
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.