CtrlK
BlogDocsLog inGet started
Tessl Logo

generate-python-runnability-test

Generates a lightweight `tests/test_runnability.py` for a Python recipe. The test just imports the recipe's agent module and asserts that `root_agent is not None` (and `app is not None` if the module defines one). The skill parses agent.py with `ast` to figure out which import-time side effects need mocking (`vertexai.init`, `google.auth.default`) and which env vars need setting (`GOOGLE_CLOUD_PROJECT`, `INTEGRATION_TEST`), and only emits the boilerplate the recipe actually needs. Runs in dry-run (report + preview) and apply (write to disk) modes. Use when the user wants to "add a runnability test", "generate test_runnability.py", "create a smoke test for the recipe", or fix the missing-required-file failure from `python-validate-recipe.yml`.

72

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured operational manual with excellent command/response specification and strong safety rules. Its weaknesses are density and altitude: internal algorithm details and historical-bug rationales inflate the body, and final verification of the generated test is left to the user.

Suggestions

Trim the 'Historical bug closed by this' narratives and the ast.walk/uv-run justification paragraphs to one-line comments in the script itself, keeping only the behavior they motivate — this would lift conciseness without losing operator-relevant facts.

Move the step 1-6 detection algorithm detail (including the import_support table rationale) into a short reference file or the script's docstring, leaving SKILL.md with what the agent needs to run and interpret the tool.

Add a post-apply verification step the skill performs itself (e.g. re-run the generator or compile/import-check the generated test) instead of only reminding the user to run pytest.

DimensionReasoningScore

Conciseness

The operational sections (Input, Run, Respond) are tight, but several passages over-explain: the multi-sentence uv-run-vs-bare-python3 rationale ('the system python3 on macOS can still be an old version'), the extended ast.walk justification ('a call nested in a function body still flags the recipe; the resulting patch is a harmless no-op...'), and two multi-sentence 'Historical bug closed by this' narratives. This fits anchor 3 ('mostly efficient but includes some unnecessary explanation or could be tightened') better than anchor 4, where over-explanation would be only minor.

3 / 5

Actionability

Fully executable guidance: exact bash commands for dry-run, apply, overwrite, and --agent-file modes; a complete flag table with required/optional status; enumerated JSON output fields; a specified response format with table columns, allowed status values, and the exact closing question ('Want me to write this file?') and next-steps block. Copy-paste ready throughout.

5 / 5

Workflow Clarity

Clear sequence with strong checkpoints: always dry-run first, show generated content for user review, explicit handling for refused_overwrite (offer --overwrite) and error (surface verbatim and stop), and never silently overwrite the conftest. It falls short of anchor 5 because validation of the outcome is delegated to the user ('remind the user to run the test locally') rather than the skill verifying the generated test passes or that the import resolves after writing — anchor 5 expects feedback loops the skill itself executes.

4 / 5

Progressive Disclosure

Good structure: the script path 'scripts/generate_runnability_test.py' is a real, well-signaled one-level reference verified on disk, and sections (What This Skill Does, Edit safety, Rules for the Agent, Input, Run, Respond) are cleanly navigable. Not anchor 5: the ~180-line body inlines substantial material that reads like script-internal documentation — the detailed six-step detection algorithm, the import_support decision table rationale, and the historical-bug narratives — which belongs at a shallower summary level or in a reference file given the script is the authority.

4 / 5

Total

16

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

Excellent description: dense, concrete, third-person, with explicit trigger phrases that map to real user requests including a CI-failure-driven entry point. Slightly long but every clause states a specific capability or trigger rather than padding.

DimensionReasoningScore

Specificity

Lists multiple concrete actions: 'Generates a lightweight tests/test_runnability.py', 'imports the recipe's agent module and asserts that root_agent is not None', 'parses agent.py with ast to figure out which import-time side effects need mocking (vertexai.init, google.auth.default)', 'which env vars need setting (GOOGLE_CLOUD_PROJECT, INTEGRATION_TEST)', plus dry-run and apply modes. Coverage is comprehensive and every claim is concrete; it does not fit the anchor 4 example, which still implies minor gaps in coverage.

5 / 5

Completeness

Explicitly answers both: 'what' is a precise statement of the generated test's behavior and detection logic, and 'when' is an explicit 'Use when the user wants to...' clause with four concrete trigger phrases. Matches the anchor 5 example structure exactly.

5 / 5

Trigger Term Quality

Natural phrases users would actually say are quoted verbatim: 'add a runnability test', 'generate test_runnability.py', 'create a smoke test for the recipe', plus the CI failure trigger 'fix the missing-required-file failure from python-validate-recipe.yml'. These include synonyms (runnability test / smoke test / test_runnability.py) and a workflow-derived trigger; nothing common is missing.

5 / 5

Distinctiveness Conflict Risk

Clear niche (runnability/smoke test for Python recipes in a specific monorepo layout) with distinct triggers including a named CI workflow failure. It would not plausibly fire for unrelated test-generation or Python skills; conflict risk is minimal.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
google/adk-samples
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.