CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/cache-key-discriminator-audit

Audits whether a cache key carries every discriminator the cached response actually depends on, so two requests that must not share a slot cannot collide. Ranks identity discriminators (tenant, user, authorization scope, plan tier) above presentation ones (locale, region, currency, feature flag), maps each missing discriminator to the data-exposure or wrong-content consequence it causes, classifies each key-and-value pair into a critical / high / medium severity band (including the Python lru_cache-on-an-instance-method trap), and writes the fix as a namespaced key builder plus the matching HTTP Vary header. Use when a cache key is being designed or changed, when a per-user or per-tenant response is about to be stored in a shared cache or CDN, or when investigating a report that one user or tenant saw another's data.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable, well-structured audit skill: the derivation procedure, severity bands, fix code, and output template are all directly executable, the workflow carries explicit validation checkpoints, and the two reference files are real, one level deep, and correctly hold only citation material. The single weakness is length — the central disclosure-vs-correctness triage framing is repeated across four sections and the exclusions section over-elaborates what could be a two-line heuristic.

Suggestions

State the identity-gaps-are-disclosures / presentation-gaps-are-correctness triage rule once (the intro is the natural home) and cut its re-statements in the discriminator-table preamble, derivation step 6, and the exclusions section.

Compress 'What this audit does not cover' to the three bullet labels plus the closing symptom heuristic ('If the data is old → invalidation/TTL; if the data belongs to someone else → this audit'); the bullet elaborations largely restate content already covered.

Trim the lru_cache section to the two-class collision distinction and the ordered three-part fix, moving the retention mechanism and citations entirely into references/lru-cache-trap.md, which already covers them.

DimensionReasoningScore

Conciseness

The body is dense and mostly non-obvious (discriminator table, severity bands, Vary gotchas, the lru_cache mechanism), but the core identity-gaps-are-disclosures / presentation-gaps-are-correctness framing is restated at least four times — in the intro ("Triage it as a security defect, not as a performance or freshness issue"), the discriminator-table preamble ("Identity and authorization discriminators come first... their absence produces wrong content, which is a correctness defect and not a disclosure"), derivation step 6, and the exclusions section — and the opening rhetoric ("The worst version of this is not slow, and it is not stale") plus the ~20-line 'What this audit does not cover' section could be cut to a few lines. This sits between 'some unnecessary explanation or could be tightened' (3) and 'minor instances of over-explanation' (4); it is not 4 because the repetition is systematic rather than minor, and not 2 because almost nothing explains basics Claude already knows — no token is spent on what caching or HTTP is.

3 / 5

Actionability

Fully executable throughout: a complete, copy-paste key-builder function with a keyword-only mandatory tenant_id that raises; exact HTTP directives ("Cache-Control: private, max-age=300" / "Vary: Authorization, Accept-Language, X-Tenant-Id"); a before/after worked example with the literal key strings ("dashboard:1" → "s:v1:t:acme:u:1:l:en-GB:dashboard"); and an exact markdown output template down to the finding-row fields. Matches the anchor-5 example's standard of copy-paste-ready guidance covering the common cases; not 4 because no pseudocode or hand-waving appears anywhere.

5 / 5

Workflow Clarity

The derivation is a numbered six-step sequence ("Pin the response, not the endpoint" through "Check the tier") used as an explicit checklist ("Read the table as a checklist against one concrete response... Only the third answer is a finding"), with concrete validation checkpoints: the conditional downgrade rule ("Downgrade to informational only after confirming identifiers are globally unique, such as UUIDs"), the live verification instruction ("Verify with a live two-principal request against the edge"), and the hard finish condition ("A key fix with no test that exercises two distinct principals against the same slot is not finished"). This is a read-only audit so the destructive-operation cap does not apply, and the sequence, checklist, and verification feedback match the anchor-5 example; not 4 because validation is explicit at both the finding and fix stages rather than having minor gaps.

5 / 5

Progressive Disclosure

Both inline references — [references/lru-cache-trap.md] and [references/http-cache-directives.md] — exist, are exactly one level deep, and are clearly signaled at the point of need with the in-body text deliberately reduced to the decision-relevant summary ("Full mechanism and Python-doc citations", "Full RFC and OWASP citations") while verbatim quotes live in the reference files. This matches the anchor-5 structure of a clear overview with well-signaled one-level-deep references; not 4 because the split is exactly right: everything needed to act is inline, and only source material is external.

5 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: third person, dense with concrete capabilities, and closed by an explicit 'Use when' clause covering design-time, pre-storage, and incident-response triggers. Every clause states a capability or a trigger, so its length is informative rather than padded.

DimensionReasoningScore

Specificity

The description enumerates concrete, distinct actions across the full audit lifecycle: "Audits whether a cache key carries every discriminator", "Ranks identity discriminators (tenant, user, authorization scope, plan tier) above presentation ones (locale, region, currency, feature flag)", "classifies each key-and-value pair into a critical / high / medium severity band", and "writes the fix as a namespaced key builder plus the matching HTTP Vary header". This matches the anchor 'Lists multiple specific concrete actions; comprehensive coverage'; it is not 4 because there are no material gaps — audit, ranking, consequence-mapping, severity classification, and the fix are all named, including the lru_cache trap.

5 / 5

Completeness

Both halves are explicit: the 'what' is a multi-clause capability list, and the 'when' is a full "Use when..." clause with three concrete trigger scenarios (key design/change, pre-storage of a per-user/per-tenant response in a shared cache or CDN, investigating a cross-principal data report). This exceeds the anchor-5 example's level of explicitness; not 4 because the 'when' is already fully explicit rather than 'could be more specific'.

5 / 5

Trigger Term Quality

Natural trigger phrases a user would actually say are present: "a cache key is being designed or changed", "a per-user or per-tenant response is about to be stored in a shared cache or CDN", and the incident-report phrasing "one user or tenant saw another's data", alongside domain nouns (tenant, user, authorization scope, plan tier, feature flag, Vary, lru_cache). Matches 'Comprehensive coverage of natural terms including synonyms'; not 4 because no commonly used variant phrasing for this niche is missing — even the bug-report wording users would use verbatim is included.

5 / 5

Distinctiveness Conflict Risk

The niche is narrow and the triggers are its own — cache-key discriminator auditing for cross-user exposure, with named specifics (Vary header, lru_cache trap, shared cache/CDN storage). It would not fire for general caching, performance, or invalidation questions because those symptoms are explicitly distinguished in the trigger phrasing; minimal conflict risk per the anchor-5 example.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents