Archive irreplaceable raw data (sequencing runs, imaging, field data) to write-once cloud cold storage — generate checksums, verify before upload, stream a tarball to S3 Deep Archive under Object Lock, and prove the pipeline with a restore test before trusting it at scale. Use when the user asks to back up or archive raw data, set up an immutable/WORM backup, generate md5 checksums for a data folder, or verify/restore an existing archive.
77
97%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Passed
No findings from the security scan
Use this skill when the user wants irreplaceable raw data pushed to durable, tamper-resistant cloud storage — or wants to verify or restore an archive that is already there.
Keywords: backup, archive, cold storage, Deep Archive, Glacier, Object Lock, WORM, immutable, checksums, md5, tarball, restore, retrieval, raw data, sequencing run, off-site copy, 3-2-1.
nohup … < /dev/null &) with a log. A terminal
Ctrl-C or dropped SSH session kills a multi-hour upload otherwise.Read ~/.config/secure_raw_data_backup/config.sh.
STORAGE_ROOT, BUCKET, PROFILE,
STORAGE_CLASS, HASH_JOBS.references/config.example.sh, asking for each value. Never store bucket
names, account IDs, or paths in the skill repo.references/bucket_setup.md — it covers
Object Lock, versioning, the incomplete-multipart lifecycle rule, and a
least-privilege IAM policy that can write but not delete.Survey candidate folders and report a table before touching anything: folder, size, file count, checksum coverage.
for d in "$STORAGE_ROOT"/*/*/; do
n=$(find "$d" -type f ! -name '*.md5' | wc -l)
e=$(find "$d" -name '*.md5' -exec cat {} + 2>/dev/null | wc -l)
printf '%-40s %6s files=%-8s md5_entries=%s\n' "${d%/}" "$(du -sh "$d" | cut -f1)" "$n" "$e"
doneInterpreting coverage: md5_entries counts lines across all .md5 files, so
it should equal the data-file count. Shortfalls are usually harmless report
files (.html data reports, README.md, logs, metadata .csv) — check which
files are uncovered before assuming a gap matters:
find "$d" -type f ! -name '*.md5' | while read -r f; do [[ -e "$f.md5" ]] || echo "$f"; doneAsk the user which folders are in scope. Data that is derived, reproducible, or already public (reference genomes, downloaded databases, assembled results) is usually not worth cold-storage money — only raw, irreplaceable data is.
Freeze the folder read-only BEFORE hashing it. This ordering is the whole ballgame, and getting it backwards is the most common way an archive goes wrong:
chmod -R a-w "$dir" # then hash; unfreeze only to write the checksum fileChecksums taken while a folder is still being worked in go stale silently.
Worse, md5sum --check only validates the files it lists — files added after
hashing pass unnoticed, so a stale checksum file is quietly incomplete as well
as wrong. A real instance: a folder was hashed mid-analysis, and by upload time
a tooling settings file had changed and a derived FASTQ had been regenerated at
double its size, while three new files had appeared and one had been deleted.
The upload guard caught the changed file; nothing would have caught the added
ones.
If the folder is still active, say so and offer the choice explicitly: wait until the work concludes, archive only the immutable raw inputs, or take a deliberate mid-work snapshot. Do not quietly archive a moving target.
scripts/generate_checksums.sh <folder> [<folder> ...]
Writes one aggregate checksums.md5 at each folder root, with paths relative
to that root (./sub/dir/file), then chmod 444. Skips folders that already
have one.
Why it is built the way it is — these are correctness requirements, not style:
md5sum --check run from that root verifies
everything in one pass. When a folder already uses per-file sidecars, match
that convention for the few missing files instead of mixing styles.Two things to tell the user explicitly:
If a folder is not writable (owned by another user), say so and stop for that
folder; do not chmod someone else's data.
Default: uncompressed tar. Justify it if asked, and re-measure rather
than assuming:
f=<one representative file>; ls -l "$f"; gzip -c "$f" | wc -cOn typical raw data (.fastq.gz, .pod5, .bam, .cram, .jpg, .zip)
compression recovers a fraction of a percent — measured 0.004% on a real
425 MB fastq.gz. Against that, tar.gz costs:
.tar.gz must be fetched and decompressed whole.Compression is worth it only for genuinely uncompressed input — raw .bcl
runs, .sam, .vcf, .fasta, XML/log trees. Check before deciding.
scripts/backup_to_cloud.sh [--dry-run] [<folder> ...]
Always start with --dry-run and show the user the plan. Then run detached:
cd "$STORAGE_ROOT" && nohup scripts/backup_to_cloud.sh > backup_$(date -Idate).log 2>&1 < /dev/null &What the script guarantees:
*.md5 in the folder and refuses to upload on any
mismatch; warns loudly if none exist.tar -cf - straight to S3 — no local temp copy, so no scratch space
is needed for a 1 TB folder. --expected-size is passed (overestimated) so
multipart part-sizing never runs short on a large archive.yes = always, no = never, s = ask again).<key>.meta.txt (ingest checksum, size, file count) and then
<key>.manifest.txt (every path and byte size) to Standard, so the
archive's contents are readable instantly without paying for a restore. The
manifest is written last and acts as the completion marker, so an interrupted
upload is retried on the next run rather than being mistaken for done.If a previous run died, check for orphaned multipart uploads before retrying —
they are billed as storage while invisible in ls:
aws --profile "$PROFILE" s3api list-multipart-uploads --bucket "$BUCKET"For each uploaded folder:
aws s3api head-object --bucket "$BUCKET" --key "<key>" --checksum-mode ENABLED
— confirm size is slightly over the folder size (tar headers and padding),
storage class is as intended, and record the ingest checksum.<key>.manifest.txt against a freshly generated on-disk listing; every
path and size must match.This validates ingest and contents. It does not prove the tar stream is extractable — only Step 6 does.
See references/verification.md for the full procedure. Summary: request a
Bulk retrieval, wait, then stream the restored object through tar -t and
compare the member list against the manifest and checksums.md5. For full
confidence on a small folder, extract to scratch and run
md5sum --check checksums.md5 against the extracted copy.
Restore is the only test that exercises the whole path. Do it before large data depends on the pipeline, and treat "the upload succeeded" as unproven until it passes.
Tell the user what is now backed up, what it costs per month, what was
excluded and why, and what remains unverified. If the local data is meant to
become read-only after archiving, chmod it — but never delete it as part of
this skill.
Quote real numbers before spending money; see references/cost_model.md. The
shape to remember: Deep Archive storage is ~$1/TB/month, retrieval is cheap,
and egress is usually the dominant cost of any verification — a 48 GB
restore-and-read is ~$0.12 Bulk retrieval plus ~$4.30 egress. Retrieval also
has latency measured in hours (Bulk up to 48 h), so plan verification around
the wait rather than blocking on it.
references/bucket_setup.md — bucket creation, Object Lock, versioning,
lifecycle rules, least-privilege IAM policy, credential profile.references/verification.md — checksum conventions, restore-test procedure,
what each check does and does not prove.references/cost_model.md — storage classes, retrieval tiers, egress, and
worked examples.references/config.example.sh — placeholder config layout.scripts/generate_checksums.sh — parallel checksum generation, NFS-safe.scripts/backup_to_cloud.sh — verify, stream tarball to cold storage, write
meta + manifest markers.5bdff99
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.