CtrlK
BlogDocsLog inGet started
Tessl Logo

llm-application-dev

Building applications with Large Language Models - prompt engineering, RAG patterns, and LLM integration. Use for AI-powered features, chatbots, or LLM-based automation.

56

Quality

65%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/llm-application-dev/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized pattern catalog with lean, mostly executable TypeScript examples covering prompts, API integration, RAG, and error handling. Its weaknesses are missing validation checkpoints in the workflows, reliance on undefined helper functions, and no progressive disclosure into reference files despite its length.

Suggestions

Add validation checkpoints and feedback loops (e.g., a concrete output-guardrail step after generation, and a validate-and-retry flow) so the RAG pipeline matches the error-handling rigor already shown.

Define or annotate the assumed helpers (embedText, vectorDb, supabase, sleep, Message, ChunkOptions) so the RAG and error-handling examples are copy-paste runnable.

Split advanced material (API patterns per provider, RAG internals) into one-level-deep reference files (e.g., references/rag.md, references/api-patterns.md) and keep SKILL.md as a lean overview.

Trim boilerplate SDK setup that Claude already knows and keep only the non-obvious parts (streaming, backoff on 429, pgvector RPC).

DimensionReasoningScore

Conciseness

The body is lean — section headers, terse comments, code, and a bullet list with no prose padding — but sections like the basic OpenAI/Anthropic client setup restate boilerplate Claude already knows, matching 'efficient; minor instances of over-explanation that could be trimmed'.

4 / 5

Actionability

Concrete TypeScript throughout (streaming generator, chunker with overlap, retry with exponential backoff), but RAG and embedding examples depend on undefined helpers (embedText, vectorDb, supabase, sleep, Message, ChunkOptions), so they are not fully copy-paste runnable — 'mostly executable; concrete code with minor gaps'.

4 / 5

Workflow Clarity

The RAG pipeline is sequenced (embed → search → build context → generate) and error handling shows retry/backoff, but there are no validation checkpoints — 'Guardrails: Validate outputs before using' is only a bullet with no workflow — matching 'steps listed but validation gaps; checkpoints missing or implicit'.

3 / 5

Progressive Disclosure

Sections are clearly headed, but ~210 lines are entirely inline with no reference files; the API patterns and RAG material would fit separate reference files, matching 'some structure; content that should be separate is inline'. The under-50-line simple-skill exception does not apply.

3 / 5

Total

14

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid description with an explicit 'Use for' clause and natural trigger terms like 'chatbots'. Its main weaknesses are topic-label phrasing instead of concrete actions, and moderate overlap risk with narrower prompt-engineering or RAG skills.

Suggestions

Rewrite the 'what' clause with concrete action verbs (e.g., 'Design prompts, build RAG pipelines, and integrate LLM APIs into applications') instead of topic labels.

Add natural trigger synonyms users would say, such as 'GPT integration', 'embeddings', 'vector search', or 'AI chatbot'.

Sharpen the 'when' clause with concrete trigger phrases (e.g., 'Use when the user wants to add an AI/chatbot feature, call OpenAI or Anthropic APIs, or build a retrieval pipeline') to reduce overlap with adjacent skills.

DimensionReasoningScore

Specificity

The description names the domain ('Building applications with Large Language Models') and three areas ('prompt engineering, RAG patterns, and LLM integration'), but these read as topic labels rather than the multiple concrete actions of a 4 or 5; it is not comprehensive.

3 / 5

Completeness

It has a clear 'what' and an explicit 'Use for...' trigger clause; the 'when' is present but could be more specific with concrete trigger phrases, matching the 4 anchor rather than the fully explicit 5.

4 / 5

Trigger Term Quality

'AI-powered features, chatbots, or LLM-based automation' are phrases users would naturally say, giving good keyword coverage, though common synonyms like 'GPT', 'embeddings', or 'AI app' are missing.

4 / 5

Distinctiveness Conflict Risk

'LLM-based automation' and 'AI-powered features' are broad enough to overlap with dedicated prompt-engineering, RAG, or agent-building skills, so it could still trigger for the wrong skill despite being somewhat specific.

3 / 5

Total

14

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
MoizIbnYousaf/ai-agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.