CtrlK
BlogDocsLog inGet started
Tessl Logo

testing-and-ci

Choose and run unit, E2E, Caddyfile fixture, automation, and CI checks for caddy-security. Use for coverage requirements, test evidence, failures, and report workflows; official OP conformance remains opt-in.

68

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable operational skill: every workflow is backed by exact commands, file paths, and validation/feedback loops, and the heavy detail is properly split into one-level-deep reference files. The main cost is conciseness — the inline CI narrative and a large set of pinned version numbers inflate the body and will drift stale. Moving runner/workflow minutiae into a reference and pointing to authoritative sources for current versions would lift both conciseness and progressive disclosure.

Suggestions

Move the GitHub-workflow narrative detail (runner/toolchain versions, TEST_MEMORY_MB, retention and permission specifics) into a reference file such as references/ci-workflow.md, keeping the body to the make targets, scope boundaries, and the link to the workflow — this serves both conciseness and progressive_disclosure.

Replace inline pinned version numbers (Go 1.26.8, tested v1.1.0, Ubuntu 24.04, Node 24, 4,200 seconds, 6144 MB) with pointers to their authoritative sources (build.yml, go.mod, Makefile defaults) so stale numbers cannot mislead a future run.

Tighten the lifecycleAddress and guard paragraphs to rule + one-line reason, keeping the QUIC-port and process-tree rationale in the referenced test-resources doc instead of restating it.

DimensionReasoningScore

Conciseness

The body is dense with genuinely non-obvious repo facts (lifecycleAddress QUIC rationale, tested contract, guard behavior) and explains nothing Claude already knows, but at 313 lines it could be tightened: the CI Workflow section narrates runner/toolchain detail ('Ubuntu 24.04 and Go 1.26.8 with GOTOOLCHAIN=local, plus Node 24, Python 3 and NSS utilities', '4,200 seconds', 'TEST_MEMORY_MB=6144 (6 GiB)', '14-day retention') that is both stale-prone and partly operational noise for the skill's core purpose. This fits anchor 3 (mostly efficient but could be tightened) rather than 4, because the volume of inline pinned configuration exceeds 'minor instances'.

3 / 5

Actionability

Commands are copy-paste ready and cover the common cases: 'go test -run TestParseCaddyfileAuthorization ./...', 'make test TEST=... TEST_DIR=.', 'make qtest', 'make ci-check', 'make fmtcfg', plus exact file paths ('testdata/caddyfile_adapt/', '<prefix>_resolved.json'), test names, env vars, and a concrete generated-artifact list. Fully executable with no pseudocode.

5 / 5

Workflow Clarity

Multi-step processes are explicitly sequenced with validation checkpoints and error-recovery loops: 'After fixture edits, run the focused test first, then a broader command' with the exact command ladder; 'If *_tmp_input.json or *_tmp_output.json files remain after a failing runtime resolution test, inspect them, fix the fixture or implementation, and remove...'; 'If CI times out, inspect the captured test events...'; and an acceptance-criteria checklist closing the skill. Validation is pervasive, so no batch/destructive cap applies.

5 / 5

Progressive Disclosure

Both bundle references are real, one level deep, and well signaled with when-to-read conditions ('Read [test surfaces and qualification evidence](references/test-surfaces.md) when selecting existing unit/E2E tests...'), and the '#subprocess-coverage' anchor resolves in references/test-surfaces.md. It falls short of anchor 5 only because ~60 lines of GitHub-workflow narrative (runner selection, memory budget, retention/permissions) sit inline where the skill's own pattern would split such exhaustive detail into a reference file.

4 / 5

Total

17

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An excellent description: it concisely states concrete capabilities, gives an explicit 'Use for...' trigger clause with natural terms, and even delineates an opt-in boundary to avoid over-triggering. The only refinement space is slightly broader verb diversity and a few more natural synonyms (test suite, run the tests) in the trigger list.

DimensionReasoningScore

Specificity

'Choose and run unit, E2E, Caddyfile fixture, automation, and CI checks for caddy-security' names the domain plus several concrete check categories, matching anchor 4 (several specific actions, minor gaps). It is below anchor 5 because the action verbs are limited to 'choose and run' — unlike the anchor-5 pattern of multiple distinct concrete actions — and 'automation' is comparatively generic.

4 / 5

Completeness

It explicitly answers both questions with concrete trigger phrases: the what ('Choose and run unit, E2E, Caddyfile fixture, automation, and CI checks for caddy-security') and the when ('Use for coverage requirements, test evidence, failures, and report workflows'), plus a scope boundary ('official OP conformance remains opt-in'). This matches the anchor-5 example pattern exactly, not merely the both-present-but-looser anchor 4.

5 / 5

Trigger Term Quality

Natural trigger terms users would actually say are present: 'unit', 'E2E', 'CI checks', 'coverage requirements', 'test evidence', 'failures', and 'report workflows' — good keyword coverage per anchor 4. It misses anchor 5's comprehensive synonym coverage (e.g., 'run the tests', 'test suite', 'make test', 'go test') that users in this repo would commonly phrase.

4 / 5

Distinctiveness Conflict Risk

The description is anchored to one repository ('caddy-security') with niche-specific triggers (Caddyfile fixtures, E2E, coverage evidence) and an explicit boundary carve-out ('official OP conformance remains opt-in'), giving a clear niche with minimal conflict risk — anchor 5. Neighboring anchor 4 (minor overlap with closely related skills) fits less well since the trigger terms are unambiguously test/CI-phrased.

5 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 9 suspicious

Warning

referenced_paths_exist

Referenced path issues: 2 missing, 1 deeper-than-1-level

Warning

Total

14

/

16

Passed

Repository
greenpau/caddy-security
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.