CtrlK
BlogDocsLog inGet started
Tessl Logo

karpathy-guidelines

Behavioral guidelines to reduce common LLM coding mistakes. Use when writing, reviewing, or refactoring code to avoid overcomplication, make surgical changes, surface assumptions, and define verifiable success criteria.

89

1.14x
Quality

86%

Does it follow best practices?

Impact

92%

1.14x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is karpathy-guidelines in multica-ai/andrej-karpathy-skills

SKILL.md
Quality
Evals
Security

Quality

Content

93%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exemplary lean skill body: every section is terse, imperative, and free of padding, with concrete heuristics and task-transformation examples instead of abstract advice. The only mild gap is that the multi-step plan template in section 4 is a placeholder pattern rather than a worked example.

DimensionReasoningScore

Conciseness

The body is lean, imperative bullet guidance ('No features beyond what was asked', 'Match existing style, even if you'd do it differently') with no explanation of concepts Claude already knows and no padding; every line earns its place.

5 / 5

Actionability

For an instruction-only skill, the guidance is concretely actionable: worked task transformations ('"Fix the bug" → "Write a test that reproduces it, then make it pass"'), a reusable step→verify plan template, and testable heuristics like 'If you write 200 lines and it could be 50, rewrite it' cover the common cases.

5 / 5

Workflow Clarity

The four sections are clearly sequenced (think → simplify → surgical → execute) and section 4 supplies an explicit verification loop ('Define success criteria. Loop until verified'), but the plan template is schematic ('[Step] → verify: [check]') and there is no fully worked end-to-end sequence, leaving minor gaps.

4 / 5

Progressive Disclosure

The skill is a short, self-contained set of guidelines with well-organized sections and a single clearly signaled external link; no content is inlined that belongs in a separate file, and no bundle files are needed.

5 / 5

Total

19

/

20

Passed

Description

80%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit 'Use when' trigger clause and a clear, action-oriented statement of what the skill does. Its main weakness is breadth: the trigger surface covers almost any coding task, creating overlap risk with other coding-related skills.

Suggestions

Narrow the trigger clause to distinguish it from general code-review/refactor skills, e.g. 'Use when writing or modifying code where scope creep, over-engineering, or unnecessary refactoring is a risk, or when the user asks for minimal/surgical changes.'

Add natural user synonyms for the core behaviors, such as 'keep it simple', 'don't over-engineer', or 'minimal diff', to improve trigger-term coverage.

Name the concrete deliverables that signal the skill applies, e.g. 'when asked to state assumptions, write a step-by-step plan with verification checks, or transform vague tasks into testable goals'.

DimensionReasoningScore

Specificity

The description lists several concrete behavioral actions ('avoid overcomplication, make surgical changes, surface assumptions, and define verifiable success criteria') that map onto the skill's sections, but they remain more abstract than tangible operations, fitting 'several specific actions; minor gaps' rather than fully comprehensive concrete coverage.

4 / 5

Completeness

It explicitly answers both what ('Behavioral guidelines to reduce common LLM coding mistakes' plus the enumerated actions) and when ('Use when writing, reviewing, or refactoring code...') with concrete trigger phrases, matching the anchor for a complete description exactly.

5 / 5

Trigger Term Quality

'writing, reviewing, or refactoring code' are natural phrases users would say, but common synonyms like 'keep it simple', 'don't over-engineer', or 'minimal diff' are absent, matching 'good keyword coverage; a few natural terms missing'.

4 / 5

Distinctiveness Conflict Risk

The trigger 'writing, reviewing, or refactoring code' spans nearly all coding activity and overlaps with general code-review and refactoring skills; only the purpose clause (overcomplication, surgical changes) distinguishes it, so it 'could still overlap with similar skills'.

3 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
OpenBMB/PilotDeck
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.