Reviews and guides LLM/AI application engineering: prompt design, prompt caching, multimodal inputs, RAG, agent loops and tool design, resilience (rate limits, retries, fallbacks), memory, model migration, evals, testing, prompt-injection defence, and observability. Synthesises practices from Anthropic, OpenAI, Google, OWASP LLM Top 10, and practitioners (Hamel Husain, Eugene Yan, Chip Huyen). Triggers on "review my prompt", "design a system prompt", "optimise tokens", "set up RAG", "build an agent", "handle rate limits", "migrate to a new model", "write evals", "test my prompt", "audit AI code", "/ai-engineering".
68
85%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
Low
Low-risk findings worth noting
Prescriptive guidance for building and reviewing LLM/AI applications. Thirteen orthogonal concerns — load only the rules the current task needs.
This
SKILL.mdis a thin index. Detailed rules live inrules/*.mdand load on demand. Curated source URLs live inreferences/primary-sources.md. Date-flagged changes since 2024 live inreferences/recent-changes.md. Literal scaffolding lives intemplates/.
Parse $ARGUMENTS (first token) and detect the mode:
| Mode | Default | Trigger |
|---|---|---|
guide | yes | Default. Open question ("how should I …", "what's the best way to …"). |
review | $0 == "review", or a file/path is supplied as $ARGUMENTS. | |
design | $0 == "design", or "scaffold a prompt / system prompt / eval". |
State the detected mode and the area(s) in scope before continuing:
Mode: review
Areas: prompt-writing, system-prompt-design
Targets: src/agents/triage.ts (system prompt at L24-78)Map the user's request to one or more rule files. Load only the rules listed for the matched area(s).
| Area | Rule file | Load when |
|---|---|---|
| Writing user prompts | rules/prompt-writing.md | "improve this prompt", few-shot questions, output format, CoT, structured outputs. |
| Designing system prompts | rules/system-prompt-design.md | Persona, tool docs, ordering, refusals, agent stop conditions. |
| Token cost / latency | rules/token-optimization.md | Prompt caching, model routing, batching, streaming, max_tokens. |
| Multimodal | rules/multimodal.md | Image/audio/PDF inputs, vision-vs-OCR, voice agents, image token costs. |
| Retrieval-augmented generation | rules/rag.md | Chunking, embeddings, hybrid search, reranking, query rewriting. |
| Agents & tool use | rules/agents-and-tools.md | Tool schemas, agent loops, parallel tool calls, error recovery, workflow vs agent. |
| Resilience | rules/resilience.md | Rate limits (429), retries with jitter, circuit breakers, fallback chains, timeouts, idempotency. |
| Memory & long-running state | rules/memory-and-state.md | Conversation summarisation, structured memory, vector memory, memory tools, compaction. |
| Model migration & versioning | rules/model-migration.md | Pin snapshots vs aliases, A/B a new model, deprecations, cross-provider migration, rollback. |
| Evaluation | rules/evals.md | Golden sets, LLM-as-judge, regression CI, error analysis. |
| Testing (engineering) | rules/testing.md | Unit/integration tests, mocks, VCR cassettes, snapshot tests, CI cost discipline. |
| Safety & guardrails | rules/safety-and-guardrails.md | Prompt injection, jailbreaks, output validation, PII, scope control. |
| Observability & versioning | rules/observability-and-versioning.md | Tracing, prompts-as-code, A/B releases, rollback. |
evals.md covers product-quality measurement (golden sets, LLM-as-judge).
testing.md covers engineering-correctness tests (mocks, VCR, snapshots).
Load both when the user asks "how do I test this?" without specifying.
Observability composition.
When the area is Observability & versioning and the task involves OTEL
wiring or attribute naming, invoke the dash0 OTEL skills via Skill()
before answering — they hold the source of truth for spans and
gen_ai.* semconv:
Skill("otel-instrumentation", ...) — SDK setup, exporters, span shape.Skill("otel-semantic-conventions", ...) — attribute naming and
spec validation.If neither skill is in the available-skills list, fall back to the
inline guidance in rules/observability-and-versioning.md and the OTEL
spec.
If the user does not name an area, ask one batched clarifying question listing the thirteen options before loading rules.
guide (default)references/primary-sources.md when a claim
is non-obvious or model-version-specific.references/recent-changes.md and
call out the date in the answer.review$ARGUMENTS.path:line.Do not edit the file in review mode unless the user asks for fixes.
designtemplates/system-prompt-skeleton.md.templates/tool-description.md.templates/eval-rubric.md.templates/golden-set.md.Load on demand — do not preload.
| Area | Files |
|---|---|
| Prompt writing | rules/prompt-writing.md |
| System prompts | rules/system-prompt-design.md |
| Token cost | rules/token-optimization.md |
| Multimodal | rules/multimodal.md |
| RAG | rules/rag.md |
| Agents | rules/agents-and-tools.md |
| Resilience | rules/resilience.md |
| Memory & state | rules/memory-and-state.md |
| Model migration | rules/model-migration.md |
| Evals | rules/evals.md |
| Testing | rules/testing.md |
| Safety | rules/safety-and-guardrails.md |
| Observability | rules/observability-and-versioning.md |
| Source URLs | references/primary-sources.md |
| Date-flagged changes | references/recent-changes.md |
| System prompt template | templates/system-prompt-skeleton.md |
| Tool description template | templates/tool-description.md |
| Eval rubric template | templates/eval-rubric.md |
| Golden-set template | templates/golden-set.md |
tools → system → stable context → volatile context → user input.rules/prompt-writing.md.rules/model-migration.md.Retry-After, jitter retries, fall back across models,
key destructive tool calls for idempotency.
See rules/resilience.md.tools, tool_choice, or thinking mid-conversation — silent cache flush.Retry-After — wastes quota and triggers thundering herds.claude-sonnet-4-7) in production — auto-upgrade silently.user_id and you have a privacy bug.path:line and the rule it violates.design mode, the produced artefact uses the matching template.review mode, findings are prioritised; no edits without consent.39b3f44
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.