CtrlK
BlogDocsLog inGet started
Tessl Logo

temporal-python-testing

Test Temporal workflows with pytest, time-skipping, and mocking strategies. Covers unit testing, integration testing, replay testing, and local development setup. Use when implementing Temporal workflow tests or debugging test failures.

79

1.12x
Quality

72%

Does it follow best practices?

Impact

87%

1.12x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./tests/ext_conformance/artifacts/agents-wshobson/backend-development/skills/temporal-python-testing/SKILL.md

The canonical home for this skill is temporal-python-testing in wshobson/agents

SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body has a sound skeleton — test-type taxonomy, progressive-disclosure map, and working code patterns — but is undermined by three sets of duplicated sections, code examples that aren't self-contained, and references to resource files that are not present in the bundle. It reads as an overview doing double duty as a summary of files that can't be loaded.

Suggestions

Collapse the duplicated sections: merge "Key Testing Principles" into "Testing Philosophy", "Coverage Targets" into "When to Use This Skill", and "How to Use Resources" into "Available Resources" — this would cut roughly a third of the body.

Make the quick-start example fully runnable: define a minimal real workflow and activity instead of `YourWorkflow`/`args`/`expected` placeholders, and state the prerequisites (temporalio install, pytest-asyncio configuration) needed to execute it.

Ship the four referenced `resources/*.md` files in the bundle (or fix the paths to wherever they live) — the "File: resources/…" pointers currently resolve to nothing, breaking the progressive-disclosure structure the skill depends on.

DimensionReasoningScore

Conciseness

Mostly efficient lists and code, but three pairs of sections duplicate the same information: "Testing Philosophy" vs "Key Testing Principles", "When to Use This Skill" vs "Coverage Targets", and "Available Resources" vs "How to Use Resources". This matches the anchor for content that could be tightened, though it is not padded with explanations of concepts Claude already knows.

3 / 5

Actionability

Two real-API code examples (WorkflowEnvironment time-skipping fixture, ActivityEnvironment) give concrete executable structure with minor gaps — placeholders like `YourWorkflow`, `args`, and `expected` are undefined, and required setup (installing temporalio/pytest-asyncio configuration) is omitted. Not 5 because the examples are not copy-paste runnable; not 3 because they are genuine API usage rather than pseudocode.

4 / 5

Workflow Clarity

There is a clear decision guide for which test type to use and when to load each resource, but no sequenced multi-step workflow with validation checkpoints (e.g., how to run the tests, interpret failures, or iterate). This matches the anchor for a present but checkpoint-less sequence; the topic is not destructive/batch so the hard cap does not apply.

3 / 5

Progressive Disclosure

References are clearly signaled one level deep with per-file "When to load" triggers — good structure — but the four referenced files (`resources/unit-testing.md`, `resources/integration-testing.md`, `resources/replay-testing.md`, `resources/local-setup.md`) do not exist in the bundle, and the resource map is duplicated across two sections. Broken reference targets plus the duplication place this at the mid anchor rather than 4.

3 / 5

Total

13

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete capabilities, natural trigger terms, explicit what-and-when guidance, and a well-defined niche. The only weakness is mild repetitiveness in the action verbs ("test… testing… testing") that keeps specificity just below comprehensive.

DimensionReasoningScore

Specificity

Names concrete methods and coverage areas — "Test Temporal workflows with pytest, time-skipping, and mocking strategies" plus "unit testing, integration testing, replay testing, and local development setup" — matching the anchor for several specific actions with minor gaps. Not 5 because the verbs are repetitive (test/testing throughout) rather than a list of distinct concrete actions.

4 / 5

Completeness

Explicitly answers both questions: what ("Test Temporal workflows with pytest, time-skipping, and mocking strategies…") and when ("Use when implementing Temporal workflow tests or debugging test failures") with concrete trigger phrases, matching the top anchor.

5 / 5

Trigger Term Quality

Includes natural phrases users would say: "pytest", "mocking", "unit testing", "integration testing", "replay testing", and "debugging test failures" — good coverage per the anchor, but missing common synonyms like "determinism", "workflow history", or "test coverage" that would warrant a 5.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche — Temporal workflow testing with named tooling — so it is clearly distinguishable from generic testing or pytest skills and unlikely to trigger for the wrong skill. Minimal conflict risk.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Dicklesworthstone/pi_agent_rust
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.