CtrlK
BlogDocsLog inGet started
Tessl Logo

llamaindex

Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal support. Use for document Q&A, chatbots, knowledge retrieval, or building RAG pipelines. Best for data-centric LLM applications.

57

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./cli-tool/components/skills/ai-research/agents-llamaindex/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable — dense, mostly executable Python covering the full LlamaIndex surface — but it is significantly over budget and structured as a monolithic reference rather than a lean overview. Marketing metrics, benchmark tables, and framework comparisons pad the file, content that duplicates the provided reference files stays inline, and workflows lack explicit validation checkpoints.

Suggestions

Cut the marketing metrics block, performance-benchmarks table, and LlamaIndex-vs-LangChain comparison (or condense to two or three lines); they add tokens without actionable value and include time-sensitive version pins that will go stale.

Move the vector-store integrations, data-ingestion patterns, and multi-modal sections into the existing reference files (references/data_connectors.md, etc.), leaving one short inline example each — the ingestion section currently duplicates content already in the bundle.

Fix the non-runnable snippets (undefined 'nodes' in the custom retriever, the 'summary = response' structured-output claim, legacy 'pinecone.init') and add an explicit validation step to the RAG pipeline pattern, e.g. evaluating the first response with RelevancyEvaluator before persisting the index for production use.

DimensionReasoningScore

Conciseness

The 570-line body carries several padded sections Claude does not need: marketing metrics ('45,100+ GitHub stars', '1,715+ contributors'), a performance-benchmarks table, a LlamaIndex-vs-LangChain comparison table, and duplicated 'Use LlamaIndex when' lists. Version pins ('v0.14.7', 'claude-sonnet-4-5-20250929') are time-sensitive without any deprecated/old-patterns framing, matching the 'noticeably verbose; several unnecessary explanations or padded sections' anchor. Not 1 because the bulk is code rather than explanation of concepts Claude already knows.

2 / 5

Actionability

Most sections give executable copy-paste Python (installation, 5-line RAG example, agents, chat engines), matching 'mostly executable guidance; concrete code with minor gaps'. Not 5 because several snippets are not runnable as written: the custom retriever returns an undefined 'nodes', 'summary = response' does not produce a Pydantic model as claimed, 'pinecone.init(...)' is a legacy API, and later sections reuse an 'index'/'llm' that was never defined in scope.

4 / 5

Workflow Clarity

Sections follow a natural implicit order (install → ingest → index → query → agents) but are organized as a topic reference, not a sequenced workflow; validation checkpoints are largely absent, with the evaluation section (RelevancyEvaluator/FaithfulnessEvaluator) presented as a standalone topic rather than a step in any pipeline. This matches 'sequence present but checkpoints missing or implicit'; not 4 since no explicit validation steps are embedded in the documented pipelines.

3 / 5

Progressive Disclosure

Three real, one-level-deep reference files are clearly signaled at the end with per-file descriptions, but roughly 300+ lines of vector-store integrations, ingestion patterns (duplicating references/data_connectors.md), multi-modal RAG, and benchmark/comparison content remain inlined in SKILL.md. This matches 'some structure but could be better organized; content that should be separate is inline'; not 4 because substantial content that belongs in the existing bundle files was kept inline.

3 / 5

Total

12

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states both what the skill covers and when to use it, with concrete natural-language triggers. Its main weaknesses are minor: trigger coverage lacks synonyms and concrete file/package terms, and a couple of broad terms (chatbots, LLM applications) invite mild overlap with general agent frameworks.

DimensionReasoningScore

Specificity

Names the domain plus several concrete capabilities ('document ingestion (300+ connectors), indexing, and querying', 'vector indices, query engines, agents, and multi-modal support'), matching the 'several specific actions; minor gaps' anchor. Not 5 because the action verbs stay generic and coverage has gaps (e.g., structured extraction is only mentioned in the body).

4 / 5

Completeness

Explicitly answers both: what ('Data framework for building LLM applications with RAG. Specializes in document ingestion...') and when ('Use for document Q&A, chatbots, knowledge retrieval, or building RAG pipelines') with concrete trigger phrases, matching the anchor-5 example structure. It is clearly not 4, since the 'when' clause is explicit rather than merely present.

5 / 5

Trigger Term Quality

'document Q&A, chatbots, knowledge retrieval, or building RAG pipelines' gives good natural-phrase coverage users would actually say, matching anchor 4. Not 5 because synonyms and concrete artifact terms (file types, package names, phrases like 'search over my documents') are missing.

4 / 5

Distinctiveness Conflict Risk

A clear RAG/document-ingestion niche ('Best for data-centric LLM applications') keeps it mostly distinct, but broad triggers like 'chatbots' and 'building LLM applications' create minor overlap risk with general LLM/agent skills. Not 3 — the core niche and triggers are far more specific than 'Works with document files'.

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (570 lines); consider splitting into references/ and linking

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
davila7/claude-code-templates
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.