CtrlK
BlogDocsLog inGet started
Tessl Logo

code-refactoring-tech-debt

Identify technical debt from actual code and change history, estimate its impact, and prioritize bounded improvements with explicit assumptions.

52

Quality

57%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/code-refactoring-tech-debt/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The skill presents a well-sequenced, honest workflow (its disclaimers about hypothetical numbers and fabricated telemetry are a genuine strength), but it is held back by heavy padding: a generic debt-taxonomy Claude already knows, a boilerplate compatibility preamble, and long illustrative templates inlined where reference files should be. Tightening the body and moving templates to bundle files would substantially improve both conciseness and progressive disclosure.

Suggestions

Delete the 'Compatibility and maintenance' editorial preamble and cut the generic debt-type taxonomy (Claude already knows what duplicated code, brittle tests, and missing documentation are), keeping only the quantification requirements and any project-specific thresholds.

Move the metrics dashboard YAML, stakeholder report template, quality-gates config, and code-standards checklist into files under references/ (e.g. references/templates.md) and link to them from a lean SKILL.md overview.

Replace illustrative pseudocode (the PaymentService 'pass' stub, hypothetical trend numbers) either with concrete tool invocations for producing the metrics from a real repo, or trim them entirely and state the measurement step as an instruction.

DimensionReasoningScore

Conciseness

The ~390-line body extensively enumerates concepts Claude already knows (e.g. 'Duplicated Code — Exact duplicates (copy-paste)', 'No API documentation', 'Missing architecture diagrams', 'Brittle tests (environment-dependent)'), and opens with an irrelevant 'Compatibility and maintenance' editorial preamble ('Modified in AAS on 2026-09-05; original metadata and license notices are retained') — matching 'noticeably verbose; several unnecessary explanations or padded sections'. It is not 1 because the Requirements section does add non-obvious constraints (report unknown inputs, never fabricate telemetry) and the examples are structured rather than free-form padding, and not 3 because whole sections (the debt-type taxonomy, generic quality gates, developer code standards) could be cut with no loss of actionable signal.

2 / 5

Actionability

There is concrete guidance — specific thresholds ('cyclomatic complexity (>10)', 'Long methods (>50 lines)'), a worked cost model ('240 hours × $150/hour = $36,000'), and a phased facade/feature-flag migration — but much of it is illustrative scaffolding rather than executable instruction: the PaymentService implementation is literally 'pass', the impact and dashboard blocks are filled with hypothetical example numbers, and no actual commands or tool invocations for scanning a repo are given. This matches 'some concrete guidance but incomplete; pseudocode instead of executable code; missing key details'. It is not 4 because a practitioner cannot execute the inventory step as written (which tool produces the complexity/duplication metrics?), and not 2 because the quantification requirements and templates do give specific shape to the output.

3 / 5

Workflow Clarity

The eight numbered sections give a clear, logically ordered sequence (inventory → impact assessment → metrics → prioritized plan → implementation → prevention → communication → success metrics) with a defined Output Format, and there are explicit checkpoints: 'Report missing cost/usage inputs as unknown' and 'Stop and ask for clarification if required inputs, permissions, safety boundaries, or success criteria are missing'. This matches 'clear sequence with most checkpoints present; minor validation gaps'. It is not 5 because there is no validate-and-retry loop (e.g. verifying measured metrics against the repo before projecting ROI), and not 3 because the sequence and gating conditions are explicit rather than implicit.

4 / 5

Progressive Disclosure

The body has a reasonable section structure with clear headers, but it is a ~390-line monolith with no bundle files at all: content that clearly belongs in separate references — the metrics dashboard YAML template, the stakeholder report template, the debt-quality-gates config, the code-standards checklist — is all inlined in SKILL.md. This matches 'some structure but could be better organized; content that should be separate is inline'. It is not 4 because nothing is split out despite obvious candidates and no cross-file navigation exists, and not 2 because the document is well-sectioned and navigable rather than a wall of text with buried content.

3 / 5

Total

12

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names its actions concretely with honest scoping language ('bounded improvements', 'explicit assumptions'), but it omits any trigger guidance and relies on a single keyword ('technical debt') for discovery. Adding a 'Use when...' clause with natural synonyms would raise both completeness and trigger-term quality.

Suggestions

Add an explicit trigger clause, e.g. 'Use when the user mentions technical debt, refactoring priorities, legacy code, or wants a code-quality/maintenance audit.'

Include natural synonyms and variations users would say — 'refactoring', 'code cleanup', 'legacy code', 'code smells' — rather than relying solely on the phrase 'technical debt'.

Optionally name one more concrete capability (e.g. producing a prioritized remediation roadmap or debt metrics report) to round out the 'what' coverage.

DimensionReasoningScore

Specificity

The description lists three concrete, verifiable actions — 'Identify technical debt from actual code and change history, estimate its impact, and prioritize bounded improvements with explicit assumptions' — which matches the anchor 'lists several specific actions; minor gaps in coverage'. It falls short of 5 because coverage of the skill's full capability (remediation planning, metrics dashboards, prevention strategy) is incomplete, and exceeds 3 because more than 1-2 actions are named with qualifying detail ('from actual code and change history', 'with explicit assumptions').

4 / 5

Completeness

The 'what' is answered clearly (identify, estimate impact, prioritize), but there is no 'Use when...' clause or any equivalent trigger guidance, so per the judging guideline a missing 'when' caps completeness at 3 ('clear what but when is missing or only weakly implied'). It is not 2 because the 'what' half is concrete rather than vague, and cannot be 4 because the 'when' is entirely absent rather than merely imprecise.

3 / 5

Trigger Term Quality

'technical debt' is the one strong natural trigger phrase, with secondary terms like 'change history' and 'prioritize', but common variations users would actually say — 'refactoring', 'legacy code', 'code smells', 'code quality', 'maintenance burden' — are absent, matching the anchor 'some relevant keywords but missing common variations or synonyms'. It is not 4 because a user asking for a 'refactor prioritization' or 'code cleanup audit' would not naturally match this wording, and not 2 because the core phrase is domain-natural rather than generic.

3 / 5

Distinctiveness Conflict Risk

'Technical debt' with the qualifiers 'from actual code and change history' and 'bounded improvements' carves a mostly distinct niche with minor overlap risk against closely related skills like general code review or refactoring guides, matching 'mostly distinct; minor overlap risk'. It is not 5 because the prioritization/impact-estimation half could collide with generic code-review or quality-audit skills, and not 3 because the domain framing is more specific than the 'Works with document files' style of broad overlap.

4 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
sickn33/agentic-awesome-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.