CtrlK
BlogDocsLog inGet started
Tessl Logo

ai-engineering

Reviews and guides LLM/AI application engineering: prompt design, prompt caching, multimodal inputs, RAG, agent loops and tool design, resilience (rate limits, retries, fallbacks), memory, model migration, evals, testing, prompt-injection defence, and observability. Synthesises practices from Anthropic, OpenAI, Google, OWASP LLM Top 10, and practitioners (Hamel Husain, Eugene Yan, Chip Huyen). Triggers on "review my prompt", "design a system prompt", "optimise tokens", "set up RAG", "build an agent", "handle rate limits", "migrate to a new model", "write evals", "test my prompt", "audit AI code", "/ai-engineering".

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured advisory skill with unusually concrete mode detection, workflow formats, and routing, plus dense non-obvious principles. Its main defect is that the progressive-disclosure architecture points at bundle files (rules/ and templates/) that are not present, and a secondary one is a duplicated rule-file table costing ~20 lines.

Suggestions

Include the referenced 'rules/*.md' and 'templates/*.md' files in the skill bundle (or fix the paths) — 17 of the 19 referenced paths do not resolve, which breaks the on-demand loading design the skill depends on.

Merge the 'Area Routing' and 'Required Reading by Area' tables: they duplicate the same 13 rule-file links, so a single table with load conditions plus a short list of references/templates would save ~20 lines.

Add an inline validation checkpoint to the 'review' workflow — e.g., confirm each detected area maps to a rule file that was actually loaded before emitting findings — so verification is not only end-of-process via the Definition of Done.

DimensionReasoningScore

Conciseness

The body is a lean index that assumes Claude's competence — Core Principles and the Anti-patterns one-liners are dense, non-obvious content with no padding. Not 5: the 'Required Reading by Area' table repeats the same 13 rule-file links already given in the Area Routing table (~20 trimmable lines). Not 3: beyond that duplication, every section earns its place and nothing explains concepts Claude already knows.

4 / 5

Actionability

For an instruction-only skill the guidance is fully executable: exact mode-detection conditions ('$0 == "review"', 'a file/path is supplied as $ARGUMENTS'), a literal output block ('Mode: review / Areas: ... / Targets: src/agents/triage.ts (system prompt at L24-78)'), a concrete findings format (What/Rule/Fix with 'path:line' and a 'Top 3 fixes' list), a single batched question list for design mode, explicit template mapping, and exact Skill() invocation calls with a documented fallback. Per the scoring notes, absence of code in an instruction-only skill is not penalized when the guidance is this actionable.

5 / 5

Workflow Clarity

Three clearly sequenced numbered workflows with real checkpoints: state mode/areas before continuing, 'Do not edit the file in review mode unless the user asks', a Definition of Done checklist, a clarifying-question rule when no area is named, and a fallback when OTEL skills are unavailable. No destructive or batch operations, so the validation cap does not apply. Not 5: validation is mostly end-of-process (the DoD checklist) rather than inline per-step feedback loops — e.g., no step verifies that each detected area actually maps to a loaded rule file before findings are produced. Not 3: checkpoints are explicit, not implicit, and error paths are handled.

4 / 5

Progressive Disclosure

The design intent is exemplary — a self-described 'thin index' with one-level-deep, well-signaled references and per-area load conditions — but scored against the actual bundle, 17 of the 19 referenced paths (all 13 'rules/*.md' and all 4 'templates/*.md') do not exist; only 'references/primary-sources.md' and 'references/recent-changes.md' are present. Not 4: the gap is not minor organization — the majority of referenced files are absent, so the on-demand loading chain cannot actually be followed. Not 2: the structure itself is good, references are prominently signaled rather than buried, and nothing that belongs in separate files is inlined.

3 / 5

Total

16

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, comprehensive and concrete area coverage, and an explicit trigger clause with natural user phrasing. The only gaps are a handful of missing natural trigger terms for the safety, memory, multimodal, and observability areas, and some overlap risk with API-reference skills.

DimensionReasoningScore

Specificity

The description enumerates the full domain with concrete sub-actions: 'prompt design, prompt caching, multimodal inputs, RAG, agent loops and tool design, resilience (rate limits, retries, fallbacks), memory, model migration, evals, testing, prompt-injection defence, and observability'. This is comprehensive, multi-action coverage comparable to the anchor-5 example, with no vague filler.

5 / 5

Completeness

Explicitly answers both questions: 'what' via 'Reviews and guides LLM/AI application engineering: ...' with a full area list, and 'when' via an explicit 'Triggers on ...' clause with concrete phrases. This matches the anchor-5 example structure exactly; there is no explicit-when gap that would cap it at 3.

5 / 5

Trigger Term Quality

Eleven natural trigger phrases users would actually say ('review my prompt', 'set up RAG', 'build an agent', 'handle rate limits', 'migrate to a new model', 'write evals'), plus the '/ai-engineering' slash trigger. Not 5: natural terms for several covered areas are missing (e.g., 'prompt injection', 'add tracing', 'image inputs', 'conversation memory'). Not 3: coverage goes well beyond 'some relevant keywords' and spans most of the domain.

4 / 5

Distinctiveness Conflict Risk

A clear niche (LLM/AI application engineering practice) with distinctive triggers like 'audit AI code' and 'design a system prompt'. Not 5: the domain is broad and several triggers ('optimise tokens', 'migrate to a new model', 'handle rate limits') could plausibly activate provider-API-reference or general code-review skills. Not 3: the framing is unmistakably AI-application engineering rather than generic.

4 / 5

Total

18

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 34 missing

Warning

Total

13

/

16

Passed

Repository
mthines/agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.