CtrlK
BlogDocsLog inGet started
Tessl Logo

cookbook-audit

Audit an Anthropic Cookbook notebook based on a rubric. Use whenever a notebook review or audit is requested.

74

1.96x
Quality

61%

Does it follow best practices?

Impact

100%

1.96x

Average score across 3 eval scenarios

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.claude/skills/cookbook-audit/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill's workflow and report format are genuinely actionable, with a concrete validation script and a well-sequenced audit process. Its core weakness is redundancy: the body duplicates large portions of the referenced style guide and repeats the same checklist items multiple times, inflating token cost without adding information.

Suggestions

Cut the inlined Style Guidelines, Structural Requirements, Content Philosophy, and anti-patterns sections down to one-line pointers to `style_guide.md` (e.g. "Scoring criteria and good/bad examples: see `style_guide.md`"), keeping only the audit-specific workflow, report format, and checklist in SKILL.md.

Deduplicate the Quick Reference Checklist — "Model name defined as constant", "explanatory text before code", and "%%capture for pip installs" each appear 2-3 times across the checklist and prose; state each rule once.

Add a short failure-handling note to the workflow (e.g., what to report if `validate_notebook.py` fails to run or detect-secrets flags a credential), giving the audit an explicit feedback loop.

DimensionReasoningScore

Conciseness

The body is noticeably verbose and heavily duplicated: "Model name defined as constant at top of notebook" appears in both the Code Quality and Technical Requirements checklists; "explanatory text before code blocks" is stated in Structure & Organization, Code Presentation, Structural Requirements, and the anti-patterns; and the Style Guidelines / Structural Requirements / Content Philosophy sections restate material the skill itself defers to `style_guide.md` ("See style_guide.md for detailed templates and examples"). This fits anchor 2 (several unnecessary or padded sections); it is above anchor 1 because none of it is basic-concept filler Claude wouldn't know, but the redundancy is pervasive.

2 / 5

Actionability

The guidance is mostly executable: a concrete command (`python3 validate_notebook.py <path>`), named tooling paths (`scripts/detect-secrets/plugins.py`, `scripts/detect-secrets/.secrets.baseline`), a fully templated audit report format with per-dimension X/5 scoring, and an explicit 8-step workflow. This matches anchor 4 (mostly executable guidance with minor gaps); it misses anchor 5 because the actual per-dimension scoring criteria and the report's improvement examples are delegated to `style_guide.md` without any inline sample.

4 / 5

Workflow Clarity

The numbered 8-step workflow is clearly sequenced (read style guide → identify notebook → run automated checks → review generated markdown → manual review → score → report) with checkpoints such as asking the user for the path when missing and an automated pre-review validation pass, plus a coverage checklist. This fits anchor 4 (clear sequence, most checkpoints, minor gaps); it is not anchor 5 because there is no error-recovery guidance for failure cases (e.g., what to do if `validate_notebook.py` or detect-secrets errors out) — though as a read-only audit skill the destructive-operation cap does not apply.

4 / 5

Progressive Disclosure

No bundle files were provided alongside this SKILL.md (no references/, scripts/, or assets/ directories exist), so this scores the references as written: `style_guide.md` and the detect-secrets paths are real, one level deep, and clearly signaled. However, the body inlines ~150 lines of style-guide material (voice/tone, structural requirements, anti-patterns, content philosophy) that clearly belongs in the referenced `style_guide.md`, duplicating it. This matches anchor 3 (structure present, but content that should be separate is inline); it is not anchor 2 because sectioning and signaling are actually good.

3 / 5

Total

13

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description cleanly and explicitly states both what the skill does and when to use it, with a well-scoped niche and no first/second-person voice issues. Its main weakness is thin capability coverage — a single audit action with no mention of scoring or recommendations — and missing common trigger synonyms like Jupyter or .ipynb.

Suggestions

Expand the 'what' with the concrete outputs the skill produces, e.g. "Audit an Anthropic Cookbook notebook against a style-guide rubric, score each dimension, and produce an improvement report."

Add natural trigger variations such as "notebook review, audit, or feedback" and "Jupyter notebooks (.ipynb) in the Cookbook" so users phrasing the request differently still hit the trigger.

DimensionReasoningScore

Specificity

"Audit an Anthropic Cookbook notebook based on a rubric" names the domain (Anthropic Cookbook notebooks) and one concrete action (audit per a rubric), but stops there — the description never mentions the scoring, report, or improvement recommendations the skill actually produces. This matches anchor 3 (domain plus 1-2 concrete actions, not comprehensive); it is above anchor 2 because the action is concrete rather than generic, and below anchor 4 because no additional actions are listed.

3 / 5

Completeness

Both parts are present and explicit: 'what' — "Audit an Anthropic Cookbook notebook based on a rubric", and 'when' — "Use whenever a notebook review or audit is requested". This satisfies the rubric's requirement for an explicit 'Use when' clause, landing on anchor 4; it falls short of anchor 5 only because the 'what' side is thin (a single action) relative to the comprehensive what-plus-triggers example.

4 / 5

Trigger Term Quality

The trigger phrases "notebook review or audit" are natural things a user would say, but common variations are missing: no "Jupyter", ".ipynb", "notebooks", "evaluate", or "feedback". This fits anchor 3 (relevant keywords but missing common variations/synonyms); it is not anchor 4 because the gaps go beyond 'a few natural terms missing'.

3 / 5

Distinctiveness Conflict Risk

"Anthropic Cookbook notebook" plus "review or audit" carves out a clear niche with distinct triggers that no generic notebook, PDF, or document skill would claim. This matches anchor 5 (clear niche, minimal conflict risk); anchor 4 would require some noticeable overlap with a closely related skill, which is not the case here.

5 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

referenced_paths_exist

Referenced path issues: 2 missing, 2 deeper-than-1-level

Warning

Total

15

/

16

Passed

Repository
anthropics/claude-cookbooks
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.