CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-readiness-audit

Audit a documentation site for agent-friendliness: discovery, markdown delivery, crawlability, semantic structure, machine-readable surfaces, and content legibility. Use when asked to assess docs.docker.com or any docs site for AI/agent readiness, produce a scored report, compare with external scanners, or generate a remediation list. Triggers on: "audit docs for agent readiness", "how agent-friendly is docs.docker.com", "score our docs for AI agents", "review llms.txt / markdown / crawlability", "create an agent-readiness remediation plan".

74

Quality

92%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, well-structured procedural skill: clear sequence with genuine error-recovery loops, appropriately split bundle files, and concrete probe targets throughout. The only weaknesses are mild redundancy across sections and per-page fetch checks that specify what to verify but not how to run it.

Suggestions

Consolidate the docs-only-host / MCP-manifest guidance into one place (section 1 or the Notes) — it is currently stated three times across sections 1, 2, and Notes.

Add ready-to-run probe commands for the section 4 per-page checks (e.g., a curl invocation with 'Accept: text/markdown' and a redirect-following fetch), mirroring how the baseline script is invoked in section 2.

Tighten the intro and section 7, which both state the prefer-live-fetch-path principle, and drop or compress restatements of 'do not rely on the homepage alone' in sections 2 and 3.

DimensionReasoningScore

Conciseness

Largely lean — every section adds domain-specific judgment rules Claude would not infer on its own (e.g. 'Treat these as separate signals: negotiated markdown works / a stable direct markdown URL works / the page advertises the correct markdown URL'). It falls short of 5 because of repetition: the docs-only-host caveat appears in section 1, again in section 2, and again in Notes ('Do not fail a docs host for lacking MCP or plugin manifests'), and the live-fetch-over-source-tree principle is stated in both the intro and section 7. A 3 would require noticeable padding or explanation of concepts Claude already knows, which is not the case.

4 / 5

Actionability

Concrete throughout: exact paths to probe (/llms.txt, /robots.txt, /.well-known/ai-plugin.json), copy-paste-ready bash invocations of the bundled script including the CHECK_TOOL_MANIFESTS=0 variant, a hard sample floor ('Sample at least 12 pages'), enumerated page types, and a P0/P1/P2 remediation format. It stops short of 5 because the per-page fetch checks in section 4 name what to verify ('Accept: text/markdown behavior', 'redirect chain length and canonical URL consistency') without providing runnable commands or expected output shapes, and criteria like 'obvious chrome/noise' and 'closely enough for retrieval parity' are left as qualitative judgment. A 3 would mean pseudocode or high-level hints only, which underrates the specificity present.

4 / 5

Workflow Clarity

The nine numbered sections form a coherent pipeline (scope → sitewide signals → sampling → per-page checks → judgment → scoring → external comparison → remediation → report), with explicit validation and recovery loops: 'If the sitemap is missing or unusable, discover pages through internal links and note the lower confidence', and the scanner-disagreement rule 'trust the live fetch, report the mismatch explicitly'. Scoring integrity rules ('score only what you verified', 'mark non-applicable checks as N/A', 'normalize the final score against applicable points only') act as checkpoints, and the rubric's foundational caps prevent averaging away weak segments. This matches the 5 anchor's sequence-with-feedback-loops-and-checklists; the read-only audit means the destructive-operation cap does not apply.

5 / 5

Progressive Disclosure

The body is a navigable overview with clearly headed sections, and heavy material is correctly externalized: the scoring rubric lives in references/rubric.md and the output format in references/report-template.md, both linked at their point of use in sections 6 and 9, and both verified to exist with no nested references (one level deep). The bundled script is invoked by path and exists at scripts/baseline-probes.sh. This is the 5 anchor's 'clear overview with well-signaled one-level-deep references'; a 4 would require organization gaps or buried references, and none are present.

5 / 5

Total

18

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete capability enumeration, an explicit 'Use when' clause reinforced with a natural-language trigger list, third-person voice, and a distinct niche. It reads like the rubric's own good examples.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'Audit a documentation site for agent-friendliness: discovery, markdown delivery, crawlability, semantic structure, machine-readable surfaces, and content legibility', 'produce a scored report, compare with external scanners, or generate a remediation list' — covering the workflow comprehensively. This matches the 5 anchor ('multiple specific concrete actions; comprehensive coverage') and exceeds the 4 anchor, which expects minor coverage gaps; none are evident.

5 / 5

Completeness

Both halves are explicit: the 'what' enumerates the audit dimensions and outputs, and the 'when' is a dedicated 'Use when asked to...' clause followed by concrete trigger phrases. This is the 5 anchor verbatim in structure; a 4 would mean the 'when' clause were present but less specific.

5 / 5

Trigger Term Quality

The 'Triggers on:' line gives natural user phrasings — "audit docs for agent readiness", "how agent-friendly is docs.docker.com", "score our docs for AI agents", "review llms.txt / markdown / crawlability", "create an agent-readiness remediation plan" — plus synonyms (agent-friendly, AI/agent readiness, agent-readiness) and technical surface terms (llms.txt, crawlability). Comprehensive natural-term coverage; a 4 would require visible missing variants, and none stand out.

5 / 5

Distinctiveness Conflict Risk

It carves a clear niche — agent/AI-readiness auditing of documentation sites — with triggers naming llms.txt, crawlability, and agent-friendliness that no adjacent skill (general web review, SEO audit) would claim. Minimal overlap risk; not the 4 anchor, which expects some overlap with closely related skills.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
docker/docs
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.