End-to-end orchestration to prepare or update a Python recipe under core/python/, contrib/, or skills/python/ so it passes every check in .github/workflows/python-validate-recipe.yml. Runs seven phases in order on an already-in-place recipe: manifest.yaml generation, environment-variable extraction, pyproject.toml alignment, ruff format+check, per-recipe `uv lock`, runnability-test generation, and a final `py_compile` verification of the generated test file. Assumes the user has already done the manual prep (deactivated any venv, `git pull` and `uv sync` from the repo root, placed the recipe at its target path, renamed if needed). Delegates to the existing sub-skills (generate-manifest, extract-python-environment-variables, align-recipe-pyproject, generate-python-runnability-test) so the master never duplicates their logic. Pauses at fixed decision points (description mismatch, existing test regeneration) AND is free to interrupt for clarification any time a phase's output looks ambiguous, unexpected, or would benefit from a human judgment call — this is an interactive skill by design. Use when the user wants to "prepare a recipe", "update a recipe end to end", "run all the checks and fixes", "make this recipe PR-ready", or invokes it by name.
Master orchestration skill. Runs the other Python-recipe skills in the right order, with the right inputs, in a single pipeline. Use when the user wants a recipe brought fully up to standard in one go.
This is an interactive skill. It's expected to pause and ask questions when doing so genuinely improves the outcome — not just at the fixed checkpoints below, but any time a phase's output is ambiguous, surprising, or would benefit from a judgment call. See rule 5 (fixed checkpoints) and rule 6 (judgment-based interruptions) for the difference.
Two ownership placeholders are written by generate-manifest and enforced by tools/validate_manifest.py. They must be the EXACT strings below — never invent, translate, or rephrase them:
OWNERSHIP_TEAM_PLACEHOLDER = "TODO: Replace with your team name"
OWNERSHIP_POC_PLACEHOLDER = "TODO: Replace with your GitHub user ID"generate-manifest is the single source of truth for these values. This skill NEVER replaces them mid-pipeline — they are intentionally left in place so CI validation fails until a human fills them in. Replacing them lives in the user's post-pipeline TODO list (see the summary's "What you still need to do" section).
The skill assumes the user has already:
origin (git pull at the repo root).uv sync at the repo root).core/python/<name>/, contrib/<name>/, or skills/python/<name>/.git diff shows what the skill changed.If the user has NOT done these and asks you to run the skill anyway, tell them to complete the prerequisites first and stop. Do NOT run git pull, git commit, deactivate their venv, or move/rename directories on their behalf — those are deliberately out of scope.
Runs seven ordered phases against a target recipe. Each phase either invokes an existing sub-skill (or its underlying script) or runs a repo-standard command:
manifest.yaml if missing. Ownership placeholders (ownership.team, ownership.poc) are LEFT AS-IS — never replaced mid-pipeline. See "Canonical placeholder strings" above..env.example; ensure load_dotenv() is bootstrapped and python-dotenv is a dep.[tool.ruff*], raise requires-python floor, ensure [project].name matches folder, reconcile description with manifest, and ensure [[tool.uv.index]] declares public PyPI as default (needed to bypass corp Airlock).ruff format + ruff check --fix on the recipe (from the repo root, so the root ruff config wins). Must run AFTER Phase 3 — align removes any recipe-local [tool.ruff*] block, and that removal is what makes the root config the effective one. Running lint before align would check against the recipe's (often more permissive) local config and miss violations that CI will later catch.uv lock — regenerate uv.lock so it reflects the post-align pyproject.toml. Does NOT install into .venv/ — that's a heavier step the user runs after they've reviewed the diff. uv lock just resolves and records; uv sync would download and install every wheel, which is scope-creep for a "prepare" pipeline.tests/test_runnability.py if missing (or ask before overwriting).uv run --no-project python3 -m py_compile <RECIPE>/tests/test_runnability.py — a lightweight sanity check that the generated (or existing) test file is syntactically valid Python. Deterministic; does NOT execute the test, resolve imports, or require .env to exist. If it fails, the master reports the error verbatim and moves on (the summary marks Phase 7 as failed). The master does NOT attempt to diagnose or fix — that's a human review task.At the end, print a summary table and remind the user to git diff and commit — the skill never commits.
Ask for --recipe-dir up front if the user hasn't given one. All seven phases operate on the same recipe.
Confirm before starting. The pipeline touches many files. Show the user the plan (the seven phases + the target recipe path) and ask for a single "yes, go ahead" before Phase 1. Do NOT prompt again for each phase unless a decision is required (see rules 5 and 6).
Invoke sub-skill SCRIPTS directly (not the sub-skills' own agent-facing SKILL.md). Reason: sub-skills each have their own "want me to apply?" prompt. In master-orchestration mode the user has already opted into apply for the whole pipeline; individual prompts would be noise. Command lines for each sub-script are given in each phase below.
Exception for pure-instructions skills: generate-manifest has no script — it's a pure-instructions skill. For that one only, load its SKILL.md (via the skill tool) and follow it inline.
Fixed checkpoints — always pause here:
needs_input for description-matches-manifest, show both sides and ask the user to pick pyproject, manifest, or delete.tests/test_runnability.py already exists, ask whether to regenerate (default: keep existing). Regeneration uses --overwrite.agent.py was found, surface the message and offer to re-run with --agent-file <path> once the user says where the entry point is. This is the one error case with a defined recovery instead of a hard stop.error — surface the message, stop the pipeline, do NOT retry (the entry-point case above is the only exception).Judgment-based interruptions — pause when it genuinely helps. This skill is interactive by design. Beyond the fixed checkpoints in rule 5, feel free to interrupt any time doing so meaningfully improves the outcome. Some situations where a pause is appropriate:
has_root_agent: false — a legit recipe should always have one; something may be wrong).architecture.agent = "multi" where you counted only one agent, or vice versa.requires-python drops support for a version the recipe's README claims to support.uv lock in Phase 5 records a suspicious dependency in uv.lock (say, a package that renames or shadows a well-known one — reviewable via git diff uv.lock).agent.py files, no app/ package, etc.) and you're unsure which to use.Do NOT interrupt for:
When you interrupt, present the specific concern, show the relevant data, and offer concrete options — don't just say "does this look OK?".
Halt on hard error. If any phase's script exits with a non-zero code that isn't refused_overwrite (handled by rule 5), stop. Print the phase name, the error, and what's already been done. Do NOT continue past a hard error.
Carve-out for Phase 3 (align). The align script exits 1 whenever any check is left in report_only status (e.g. missing [build-system], a non-PyPI default index) — those are deferrals-by-design, NOT hard errors (see Phase 3b). Do NOT halt on Phase 3's exit code alone: decide from its JSON. Halt only if a check has status error (or an unexpected needs_input / would_fix); a non-zero exit whose only non-clean checks are report_only means "continue".
Report progress compactly. After each phase, one line: Phase N (<name>): <one-line outcome>. Do NOT dump raw JSON. Do NOT re-render each sub-skill's own table — the summary at the end covers it. Judgment-based interruptions from rule 6 are separate from progress lines and should be their own turn (question, then wait for the answer).
Never commit. The skill is done when the summary is printed. Let the user git diff and commit.
| Field | Required | Description |
|---|---|---|
| Recipe directory | Yes | Path to the recipe root (e.g. core/python/cross-session-memory, contrib/my-recipe, skills/python/my-skill). Passed to every sub-script as --recipe-dir. |
If the user has not specified the recipe directory, ask for it before proceeding.
Step 0a — Verify the recipe directory actually exists. A path typo shouldn't cost the user the whole plan-confirmation round-trip only to fail in Phase 1:
[ -d <RECIPE_DIR> ] || { echo "Recipe directory not found: <RECIPE_DIR>"; exit 1; }If it isn't a directory, stop immediately with that message — do NOT show the plan or prompt.
Step 0b — Verify the recipe folder name matches CI's naming rules. python-validate-recipe.yml's Check 1 (folder-name regex + max length) rejects folders that don't match ^[a-z][a-z-]*$ or exceed .github/policy.yml recipe_naming.max_folder_name_length. Historically the pipeline was BLIND to this — it would run all 7 phases against a folder named data_science or MyBadName, report success, and let CI reject the PR later (or worse: Phase 3's project-name-matches-folder would propagate the bad name into [project].name). This check catches it up front.
MAX_LEN=$(uv run --no-project --with pyyaml python3 .github/scripts/load_policy.py recipe_naming.max_folder_name_length)
uv run --no-project python3 .agents/skills/prepare-python-recipe/scripts/check_folder_name.py \
--recipe-dir <RECIPE_DIR> --max-length "$MAX_LEN"The check exits 0 silently on a compliant name; on violation it exits 1 with the specific offending characters, the length overrun (if any), and a suggested compliant name derived from the current one (lowercase, _ → -, drop disallowed characters, truncate on a hyphen boundary). The suggestion is ADVISORY — the script never renames anything.
If it fails, HALT the pipeline before Phase 1. Print the script's stderr verbatim (it already includes the suggestion and the manual git mv command). Do NOT show the plan, do NOT prompt to proceed, do NOT ask "want me to rename?" — renaming a recipe directory is the user's decision, not the skill's. They rename by hand and re-invoke the skill.
Only proceed past this step if the folder-name check passed.
Step 0c — Show the plan and get confirmation. Before composing the plan, glance at the recipe for anything non-standard (package not called app/, .env.example outside root, missing tests/, extra Python source dirs, deprecated model literals per AGENTS.md). If any will affect what the pipeline does, flag them briefly in the plan message so the user isn't surprised mid-pipeline. Skip the flags entirely for a standard recipe.
Then flag the assumptions the pipeline is making and show the user the plan. Do NOT frame these as "prerequisites" — they're a heads-up so the user can push back if any assumption is wrong, not a preflight checklist for the user to tick off:
A few things I'm assuming — say so if any aren't true:
- You've deactivated any active venv.
- You've run
git pullanduv syncat the repo root.<RECIPE_DIR>is already at its target path (and renamed to its final basename).I'll run the prepare-python-recipe pipeline on
<RECIPE_DIR>— 7 phases:
- Generate manifest.yaml (if missing)
- Extract env vars into .env.example
- Align pyproject.toml
- Ruff format + check --fix
- uv lock inside the recipe (regenerates uv.lock; does NOT install .venv/)
- Generate tests/test_runnability.py (if missing)
- Compile-check the runnability test (
py_compile; reports failure but does not debug)Nothing gets committed — you'll
git diffat the end. Proceed?
Get a yes-or-no. If no, stop.
1a. Check whether manifest exists.
[ -f <RECIPE_DIR>/manifest.yaml ] && echo exists || echo missing1b. If missing — load the generate-manifest skill (via the skill tool with name="generate-manifest") and follow its instructions for this recipe. That skill writes manifest.yaml with the canonical ownership placeholders — LEAVE THEM AS-IS.
1b. If exists — skip generation. Do NOT read ownership.team / ownership.poc and do NOT prompt about them. Whatever they are (real values or the canonical placeholders), leave them untouched — the user handles ownership post-pipeline.
Progress line: Phase 1 (manifest): generated | pre-existing.
Invoke the script directly:
uv run --no-project python3 .agents/skills/extract-python-environment-variables/scripts/extract_env_vars.py \
--recipe-dir <RECIPE_DIR>(No --dry-run — master runs it in apply mode. Use uv run --no-project python3, not a bare python3: the script imports tomllib, which is stdlib only on Python 3.11+, and uv's managed interpreter guarantees that.)
The script:
.env.exampleload_dotenv() into the package __init__.py if missingpython-dotenv>=1.0.0 to [project].dependencies if missing"gemini-3.5-flash") in source files with bare os.getenv(...) calls and records the substitution in .env.exampleRead the script's stdout (structured [INFO] / [PASS] / [WARN] log lines, not a table). Extract the counts: how many vars added, whether load_dotenv was injected, whether python-dotenv was added, and how many hardcoded model names were replaced (if any).
Progress line: Phase 2 (env vars): <N> vars added to .env.example, load_dotenv <injected|already present>, python-dotenv <added|already present>[, <M> hardcoded model name(s) replaced in source].
Must run BEFORE Phase 4 (lint). Align removes any recipe-local [tool.ruff*] block — after removal the recipe uses the root pyproject.toml's ruff config. If Phase 4 (lint) ran first, it would check against whatever local config the recipe shipped with (often more permissive than root), miss violations that root's config would catch, and CI would fail on files the pipeline claimed were clean. Do NOT reorder these two phases without changing this reasoning.
Invoke the align script directly. Two-pass logic:
3a. Dry-run first to detect whether description mismatch will need user input:
uv run --no-project --with tomlkit --with 'ruamel.yaml' --with packaging \
python .agents/skills/align-recipe-pyproject/scripts/align_pyproject.py \
--recipe-dir <RECIPE_DIR> --dry-runParse the JSON. If any check has status == "needs_input" for description-matches-manifest, pause: show both pyproject_description and manifest_description from details, ask the user to pick pyproject, manifest, or delete.
3b. Apply — with the description-source flag if you got one:
uv run --no-project --with tomlkit --with 'ruamel.yaml' --with packaging \
python .agents/skills/align-recipe-pyproject/scripts/align_pyproject.py \
--recipe-dir <RECIPE_DIR> \
[--description-source=<CHOICE>]The align script exits 1 (non-zero) whenever any check is report_only — that is expected, not a hard error, so do NOT apply rule 7's halt to it. Decide from the JSON, not the exit code: if the apply run's only non-clean checks are report_only, note them in the summary (the master does NOT auto-fix these) and continue; halt only if a check has status error. Two rules can produce report_only:
build-system-present (missing [build-system] — backend choice is editorial)default-pypi-index (a default index is declared but points somewhere other than public PyPI — divergence may be intentional)Progress line: Phase 3 (align): <N> fix(es) applied; <M> report-only issue(s) left.
Run both from the repo root (not the recipe dir) so the root pyproject.toml's ruff config wins. This runs AFTER Phase 3 on purpose (see Phase 3's header comment for why):
uv run ruff format <RECIPE_DIR>
uv run ruff check --fix <RECIPE_DIR>ruff check's exit code matters here — do NOT treat every non-zero code the same way:
C901 complex-structure, PLR0912/0915 too-many-branches/statements). Do NOT stop the pipeline — note them in the summary as a Manual TODO so the user can refactor or add per-file # noqa markers after review.Note on E402 and __init__.py: Phase 2 (env-var extraction) already suppresses E402 on any trailing relative import (from . import agent) in the recipe's package __init__.py — the canonical ADK-recipe pattern where env-bootstrap side effects (load_dotenv(), os.environ.setdefault(...)) intentionally precede a from . import ... line so env vars are populated before agent submodules load. If you see E402 on an __init__.py in Phase 4's output, something went wrong upstream (Phase 2 didn't detect the pattern, or a new file appeared between Phases 2 and 4). Note that in the progress line so it isn't quietly buried.
Progress line: Phase 4 (lint): <N> file(s) formatted, <M> issue(s) auto-fixed, <K> unfixable issue(s) left.
uv lockNow that pyproject.toml is stable (Phase 3 aligned it, Phase 4 didn't touch it), regenerate the lockfile so it matches:
uv lock --python 3.11Run this WITH workdir = <RECIPE_DIR> (do not cd — pass the working directory via the tool call).
Why --python 3.11 explicitly? CI's .github/workflows/python-dependency-policy.yml pins Python 3.11 when it runs uv lock --check on every recipe. If we lock here with whatever interpreter the user happens to have installed (typically 3.12 or newer on modern machines), the recipe locks cleanly locally but the CI check fails on the PR with a confusing The requested interpreter resolved to Python 3.11.15, which is incompatible with the project's Python requirement: >=3.12 — mis-reported by the workflow as "lockfile is out of date". Forcing 3.11 here surfaces the same incompatibility at pipeline time, when the user can fix it or push back, rather than at PR time. Phase 3's python-version-floor check should have already caught this and rewritten requires-python, so the lock should succeed — but pinning defends against edge cases where the check was too permissive or the recipe had a compatible-release ceiling the rewrite couldn't lower.
Why uv lock and not uv sync? The pipeline's job is to prepare the recipe, not to install and validate its runtime environment. uv lock resolves dependencies against the aligned pyproject.toml and writes uv.lock — that's the artefact CI and downstream consumers need. uv sync would additionally download every wheel into .venv/, which:
The user gets a real .venv/ by running uv sync themselves after reviewing the diff — see the "Next steps" block at the end of the summary.
If uv lock --python 3.11 fails, halt and surface the error verbatim (rule 7). The most common cause is a requires-python specifier that excludes 3.11 (>=3.12, ~=3.12, etc.) that Phase 3 could not rewrite — the fix is to either lower the recipe's floor to >=3.11 or, if the recipe genuinely needs newer features, raise the issue with the repo maintainers to update CI's pinned interpreter.
Progress line: Phase 5 (recipe lock): uv lock --python 3.11 completed.
6a. Check whether the test exists.
[ -f <RECIPE_DIR>/tests/test_runnability.py ] && echo exists || echo missing6b. If missing — generate:
uv run --no-project python3 .agents/skills/generate-python-runnability-test/scripts/generate_runnability_test.py \
--recipe-dir <RECIPE_DIR>(Use uv run --no-project python3, not a bare python3, so a modern managed interpreter runs it — matching this sub-skill's own SKILL.md and the rest of the pipeline.)
6b. If exists — pause and ask the user: keep existing (default) or regenerate. If they choose regenerate, re-run the same command with --overwrite appended.
If the script errors (no agent.py found), surface the message and offer to re-run with --agent-file <path> when the user tells you where the entry point is.
Progress line: Phase 6 (runnability test): generated | kept existing | regenerated.
Runs LAST. Lightweight sanity check that the generated (or existing) tests/test_runnability.py is at least syntactically valid Python. Deliberately weaker than uv run pytest: it does NOT execute the test, resolve imports, or require .env to be populated. Its only purpose is to catch generator bugs (invalid Python emitted by Phase 6) and gross syntax errors in a hand-edited test file.
7a. Check the test file exists. If Phase 6 skipped generation (agent.py not found, so no test was written) or the user chose not to regenerate an existing broken test, there may be nothing to compile. Skip and record it.
[ -f <RECIPE_DIR>/tests/test_runnability.py ] && echo exists || echo missing7b. Compile.
uv run --no-project python3 -m py_compile <RECIPE_DIR>/tests/test_runnability.pyUse uv run --no-project python3 here too — not a bare python/python3. Two reasons: python may not be on PATH at all on some systems, and (more importantly) a guarded test with multiple patches emits a parenthesized with (...) block, which is Python 3.10+ syntax. Compiling it under an older system interpreter would report a spurious SyntaxError on a file that is actually valid. uv's managed interpreter is 3.11+, so this is a true syntax check rather than a version artifact.
Report per outcome:
Phase 7 (verify): compile OK.Phase 7 (verify): compile FAILED — <one-line snippet of the error>.Phase 7 (verify): skipped (no tests/test_runnability.py to check).Note: passing Phase 7 does NOT mean the recipe actually runs — it means the test file is valid Python. Actually running the test (which validates that agent.py imports and root_agent is non-None) is still a manual uv run pytest step listed under "Next steps" in the summary.
While the pipeline runs, print a short progress line per phase (see above). Do NOT dump raw JSON. Do NOT re-render sub-skill tables.
Track three things as the pipeline runs so you can report them at the end:
manifest.yaml; Phase 2 → .env.example, package __init__.py, pyproject.toml, any source files where a hardcoded model name was replaced; Phase 3 → pyproject.toml; Phase 4 → any .py under the recipe; Phase 5 → uv.lock; Phase 6 → tests/test_runnability.py; Phase 7 → nothing).report_only).At the end, print the sections below in order. Section 3 is conditional — omit it entirely if nothing belongs in it. Sections 1, 2, and 4 always print.
| Phase | Outcome | Notes |
|---|---|---|
| 1. Manifest | ok | generated (ownership placeholders left for user to fill in) |
| 2. Env vars | ok | 3 added, load_dotenv injected, python-dotenv added |
| 3. Align | ok | 2 fixes applied |
| 4. Lint | ok | 12 files formatted, 4 issues auto-fixed |
| 5. Recipe lock | ok | done |
| 6. Runnability test | ok | generated |
| 7. Verify (compile-check) | ok | tests/test_runnability.py compiles |
Use plain words in the Outcome column (ok / skipped / failed). No emoji unless the user asked for them.
Short bullet list, grouped Created and Modified, one line per file. Aggregate large groups (e.g. "12 .py files formatted (Phase 4)") rather than listing each individually. Omit files that were checked but untouched. If a group is empty, drop its heading. If nothing changed at all, print Nothing changed — the recipe was already fully aligned.
Example:
Created
manifest.yaml(Phase 1)tests/test_runnability.py(Phase 6)uv.lock(Phase 5)Modified
pyproject.toml(Phases 2, 3 — addedpython-dotenv; aligned rules)<pkg>/__init__.py(Phase 2 — addedload_dotenv()).env.example(Phase 2 — 3 vars added)agent.py(Phase 2 — replaced hardcoded model name withos.getenv("MODEL_NAME"))- 12
.pyfiles formatted, 4 auto-fixed (Phase 4)
Only print this section if at least one item belongs in it. If the pipeline finished every attempted action cleanly, omit the section entirely — do NOT print a "nothing failed" placeholder. The summary table's Outcome column already conveys clean runs.
One line per item. This is for automation attempts that couldn't finish — not for deferrals-by-design (those go in Section 4).
Cases that go here:
ruff check --fix ran but couldn't auto-fix some violations. Show <file>:<line> — <codes> for each.py_compile on tests/test_runnability.py returned an error. Show a one-line snippet of the stderr.Example:
- Phase 4 (ruff) — 2 violations remain that
ruff check --fixcan't auto-fix:app/deploy.py:276(C901,PLR0915),app/tools.py:52(C901). Refactor or add# noqa: <codes>at the def line.- Phase 7 (verify) —
uv run --no-project python3 -m py_compile tests/test_runnability.pyfailed:SyntaxError: invalid syntax (line 14). Likely a generator bug or hand-edit; regenerate or fix before running pytest.
A single short TODO list. Keep every entry to one line. Standard items come first (always shown), then conditional items (only if the phase raised them), then a copy-pasteable command block. Do NOT invent items — only include what actually applies to this run.
Standard — always include (skip only if genuinely N/A):
<RECIPE_DIR>/manifest.yaml and replace ownership.team ("TODO: Replace with your team name") and ownership.poc ("TODO: Replace with your GitHub user ID") with real values. CI validation intentionally fails until you do. If the manifest pre-existed and already had real values, this is a no-op — glance to confirm..env.example — replace each <TODO: ...> placeholder with the real value (or delete the line if the variable isn't used).cd <RECIPE_DIR> && git diff — inspect every change the pipeline made before committing.Conditional — include ONLY if the phase raised it:
build-system: add a [build-system] block to pyproject.toml — see .agents/skills/align-recipe-pyproject/SKILL.md for hatchling / uv_build templates.pypi-index: [[tool.uv.index]] default points somewhere other than public PyPI — verify this is intentional or fix per the align skill.Commands — always show:
cd <RECIPE_DIR>
uv sync # install deps into .venv/ (Phase 5 only ran uv lock)
uv run pytest tests/test_runnability.py -v # confirm the runnability test actually passes
# commit when you're happyuv sync is what actually installs the recipe's dependencies into .venv/. The pipeline stopped at uv lock on purpose — installing is heavier and better done after you've reviewed the diff.
Then stop. Do NOT commit. End your turn.
a862ebc
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.