CtrlK
BlogDocsLog inGet started
Tessl Logo

skill-audit

Audit codebases for quality, consistency, and broken patterns — use for pre-release or tech debt review

49

Quality

54%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skill-audit/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

42%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill teaches a genuinely clear five-phase audit methodology with a good severity taxonomy and reporting format, but it is padded with near-duplicate restatements of the same phases and placeholder templates instead of one worked example, and it keeps everything inline in one long file. Tightening it to the phases plus one completed example, and splitting templates into a reference file, would address all four dimensions at once.

Suggestions

Cut the "Common Patterns" section (three near-verbatim restatements of the five phases) and "The Bottom Line"; the phase templates already convey the methodology, saving roughly 100 lines.

Replace several placeholder templates with one fully worked example — a real filled-in audit item showing location, expected, method, result, evidence, and its entry in the findings summary — so the format is executable rather than fill-in-the-blank.

Add validation checkpoints for the batch operation: instruct the agent to confirm the discovery list is complete before executing, and to re-verify any ⚠️/❌ finding before reporting it (e.g., re-run the check or cite the observed evidence twice).

Move the three audit templates and the remediation-plan formats into a `references/templates.md` file referenced from a short "Templates" section, keeping SKILL.md as a lean overview.

DimensionReasoningScore

Conciseness

The body is noticeably verbose for what it teaches: the five-phase process is restated nearly verbatim three more times in "Common Patterns" (Patterns 1–3 each re-summarize Scope/Discovery/Execute/Report/Fix), placeholder templates repeat the same fields across Phase 3, Phase 4, and the "Audit Templates" section, and "The Bottom Line" re-repeats the core principle already stated in the Overview. "Best Practices" items like "Document Everything" and the "Poor: Fix it" example explain things Claude already knows. Not 1 because there is no tutorial-style padding of background concepts; not 3 because the duplicated phase restatements and redundant template sections are genuinely removable at scale (the file could lose ~40% of its lines).

2 / 5

Actionability

The guidance names real tools ("Use Glob and Grep to find all relevant code", "Use task plan tool", "Use AskUserQuestion if needed") and gives a concrete severity taxonomy (Critical/Major/Minor) and an evidence-based reporting format. However, the bulk of the instruction is placeholder pseudocode templates — "[file:line]", "[what should happen]", "[how to verify]", "[N] items" — with no worked example showing a filled-in audit item end to end. This sits at 'some concrete guidance but incomplete; pseudocode instead of executable code', short of 4 where the guidance is mostly executable.

3 / 5

Workflow Clarity

The five phases (Scope → Discovery → Execution → Reporting → Remediation) are clearly sequenced with per-item pass/fail and evidence recording, and the checklist methodology is sound — alone that would rate 4. But this is a batch operation over many items, and there are no verification checkpoints on the audit itself: nothing tells the agent to re-verify ambiguous findings, confirm the checklist is complete before reporting coverage statistics, or handle items that fail mid-audit. The rubric's batch-operation cap on missing validation therefore holds this at 3.

3 / 5

Progressive Disclosure

The file has clear section headers and its external references ("Load `skills/blocks/engineering-method-selection.md` from the installed plugin", "see `skills/blocks/codex-host-adapter.md`") are one level deep and clearly signaled. But it is a ~560-line monolith: the Common Patterns restatements, the three audit templates, and the remediation-plan formats are exactly the content that belongs in a `references/templates.md` file loaded on demand, and the referenced plugin files cannot be verified within this bundle. This matches 'some structure but could be better organized; content that should be separate is inline'.

3 / 5

Total

11

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid, concise, third-person description that answers both what and when with natural trigger language. Its main weaknesses are the thin action list (one verb) and overlap risk with code-review skills that the body explicitly disambiguates but the description does not.

Suggestions

Add one or two concrete actions to the 'what' clause, e.g., "Audit codebases for quality, consistency, and broken patterns; builds a checklist, executes it systematically, and reports prioritized findings".

Sharpen the 'when' clause with concrete trigger phrases users would actually say, e.g., "use for pre-release checks, tech debt review, or when the user asks to sweep the whole app for broken features or inconsistencies".

Reduce conflict risk with code-review skills by signaling the systematic whole-codebase nature, e.g., "systematic checklist-driven audit of an entire codebase" rather than the generic "for quality".

DimensionReasoningScore

Specificity

"Audit codebases for quality, consistency, and broken patterns" names the domain and lists audit focus areas, but 'audit' is the only action verb and the focus list is generic rather than a set of concrete actions (no 'find issues', 'categorize by severity', 'generate remediation plan'). It matches 'Names domain and 1-2 concrete actions, but not comprehensive' — not 4 because the actions listed are adjectives for the audit, not distinct capabilities; not 2 because it does go beyond naming the domain.

3 / 5

Completeness

Both parts are present: what = "Audit codebases for quality, consistency, and broken patterns"; when = "use for pre-release or tech debt review". The 'when' clause is explicit but brief and could be more specific with concrete trigger phrases (e.g., "or when the user asks to check the entire app for broken features"), which is exactly anchor 4; anchor 5 requires fuller trigger-phrase coverage.

4 / 5

Trigger Term Quality

"audit", "codebases", "quality", "consistency", "broken patterns", "pre-release", "tech debt review" are natural phrases users would say. It falls short of anchor 5 ("comprehensive coverage of natural terms including synonyms") because common variations like "code review sweep", "find broken features", "comprehensive check", or "regression check" are absent; it beats anchor 3 since the included terms are exactly what a user would say.

4 / 5

Distinctiveness Conflict Risk

"Audit codebases for quality, consistency" overlaps heavily with generic code-review/quality skills — indeed the body itself must say "Do NOT use for: Code quality review (use skill-code-review)" to disambiguate, which the description doesn't convey. This fits 'somewhat specific but could still overlap with similar skills'; it is not 4 because the quality/consistency wording invites collision with code-review and simplify-type skills.

3 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (563 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
nyldn/claude-octopus
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.