CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/webhook-delivery-tester

The single webhook-testing home, sender AND receiver: build-an-X for webhook delivery + receiver tests per Standard Webhooks (standardwebhooks.com) - HMAC-SHA256 signature verification, retry semantics with exponential backoff + jitter, replay-window check via timestamp tolerance, ordering guarantees, dead-letter handling for permanent failures, content-type + body-encoding fidelity - plus inbound capture-and-replay hardening (runtime-signed fixtures, tampered-payload and future-timestamp rejection, key-rotation acceptance, sanitized production captures) in references/inbound-replay.md. Use when authoring tests for webhook senders OR receivers in any system (Stripe / Twilio / SendGrid / GitHub / GitLab outbound webhooks; SaaS app inbound webhooks), including payment and realtime integrations.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured workflow spine with real, verified one-level-deep references, a clear step sequence, and a useful anti-patterns/fix table. Its weaknesses are confined to minor redundancy and example code that leans on undefined test helpers and a placeholder instead of being copy-paste ready.

Suggestions

Make the sender/receiver test examples self-contained by defining the helper stubs (capture_outbound_webhook(), mock_receiver, sign_for_now(), send_webhook()) once, and replace the literal 'headers=...' placeholder in the Step 6 idempotency example with the actual headers dict used in Steps 4-5.

Tighten conciseness by collapsing the Step 9 checklist (it restates the step titles already covered) into a short 'minimum coverage' line, and dropping the References section's re-listing of links already cited inline.

Trim known-concept explanation such as 'Webhooks are sent at-least-once (vendor retries on 5xx); receiver must be idempotent' to a bare cross-reference to the idempotency-test-author skill, since the idempotency point is already carried by the test code.

DimensionReasoningScore

Conciseness

The body is mostly tables and code with little concept padding, but has trimmable redundancy: Step 9 recapitulates prior step titles, the References section re-lists links already cited inline, and it explains at-least-once semantics Claude already knows ('Webhooks are sent at-least-once (vendor retries on 5xx); receiver must be idempotent'). Not 5 — not every token earns its place; not 3 — the inefficiencies are minor, not sections of over-explanation.

4 / 5

Actionability

Step 2's sign_webhook is fully executable and the signature-scheme examples are concrete, but several tests call undefined helpers (capture_outbound_webhook(), mock_receiver, sign_for_now(), send_webhook(client, ...)) and one uses a literal 'headers=...' placeholder, so they are mostly-but-not-fully executable. Not 5 — not copy-paste ready; not 3 — the guidance is real code with only minor gaps, not pseudocode.

4 / 5

Workflow Clarity

Steps 1-9 are clearly sequenced with a Step 9 end-to-end checklist per side, an anti-patterns table mapping each failure mode to its fix, and explicit 'mark critical' escalation for replay and ordering gaps — a checklist plus feedback loop matching the top anchor. The workflow authors tests rather than performing destructive/batch operations, so the validation cap does not apply.

5 / 5

Progressive Disclosure

The core workflow spine (signing, verification, replay, idempotency, ordering tests) stays in SKILL.md while lookup data splits into real, one-level-deep bundle files — references/inbound-replay.md and references/vendor-payloads-and-retries.md — both clearly signaled inline where used and again in the References section. This matches the top anchor: clear overview, well-signaled shallow references, easy navigation.

5 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a dense, trigger-rich statement of what the skill does and when to use it, covering both sender and receiver sides with concrete capability names and vendor synonyms. Its only weakness is mild internal jargon and marketing phrasing ('single webhook-testing home', 'build-an-X') that pads without helping discovery.

DimensionReasoningScore

Specificity

Enumerates many concrete actions — 'HMAC-SHA256 signature verification, retry semantics with exponential backoff + jitter, replay-window check via timestamp tolerance, ordering guarantees, dead-letter handling... content-type + body-encoding fidelity' plus inbound 'capture-and-replay hardening' — comprehensive with no real coverage gaps, matching the top anchor rather than the 'minor gaps' level below.

5 / 5

Completeness

Both what and when are explicitly stated: the capability list answers 'what', and 'Use when authoring tests for webhook senders OR receivers in any system... including payment and realtime integrations' is an explicit when-clause with concrete triggers, exactly matching the top anchor.

5 / 5

Trigger Term Quality

Natural user phrasing is well covered: 'webhook senders OR receivers', 'Stripe / Twilio / SendGrid / GitHub / GitLab outbound webhooks', 'SaaS app inbound webhooks', 'payment and realtime integrations', 'signature verification' — vendor names act as synonyms, satisfying comprehensive natural-term coverage.

5 / 5

Distinctiveness Conflict Risk

A clear webhook-testing niche with vendor-named and sender/receiver-scoped triggers makes it unlikely to fire for the wrong skill. Mild internal jargon ('The single webhook-testing home', 'build-an-X') adds slight noise but creates no overlap risk with other skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents