CtrlK
BlogDocsLog inGet started
Tessl Logo

rag-and-embeddings

Retrieval-augmented generation and embeddings: retrieval pipeline (chunking, reranking, context assembly, grounding/citation), embedding-model selection, dense/sparse/hybrid retrieval. [EXPLICIT] Trigger: 'rag', 'retrieval augmented', 'embeddings', 'vector search', 'reranking', 'chunking', 'grounding'.

SKILL.md
Quality
Evals
Security

RAG & Embeddings

"Retrieval decides what the model sees; grounding decides whether you can trust what it says." [INFERENCE]

TL;DR

Designs retrieval-augmented generation end-to-end — chunking, embedding-model selection, dense/sparse/hybrid retrieval, reranking, context assembly, and grounding/citation — to reduce hallucination and keep answers traceable to sources. Complements ai-software-architecture (system design) and data-strategy (the corpus). [EXPLICIT]

When to use

  • Designing a RAG pipeline (ingest → chunk → embed → retrieve → rerank → generate). [EXPLICIT]
  • Choosing an embedding model + similarity metric (cosine/dot/L2). [EXPLICIT]
  • Deciding dense vs sparse vs hybrid retrieval for a corpus. [EXPLICIT]
  • Reducing hallucination via grounding, citation, and context budgeting. [EXPLICIT]
  • NOT for fine-tuning the base model (deferred) or general prompt design (use prompt katas). [INFERENCE]

Sub-capabilities (resource map)

CapabilityReference
RAG pipeline patternsreferences/rag-patterns.md
Embedding model selectionreferences/embedding-strategy.md

Procedure

  1. Chunk by semantic unit with overlap; size to the embedding context + retrieval granularity. [DOC]
  2. Embed with a model matched to domain/language; fix the similarity metric. [DOC]
  3. Retrieve top-k (hybrid dense+BM25 when recall matters), then rerank for precision. [INFERENCE]
  4. Assemble context within a token budget; dedupe; preserve provenance. [INFERENCE]
  5. Ground the answer: cite sources, refuse when retrieval is empty/low-confidence. [DOC]

Quality Criteria

  • Chunking preserves semantic units with controlled overlap. [INFERENCE]
  • Embedding model + metric chosen for domain, not by default. [DOC]
  • Retrieval + rerank balance recall and precision. [INFERENCE]
  • Answers cite sources; empty/low-confidence retrieval triggers refusal. [DOC]
  • Claims evidence-tagged. [EXPLICIT]

Anti-Patterns

  • Fixed-size character chunking that splits mid-sentence/table. [INFERENCE]
  • Single-vector dense retrieval where exact-term recall matters (use hybrid). [INFERENCE]
  • Generating without citation or refusal path (silent hallucination). [DOC]

Contract

  • Aceptación: pipeline ingest→chunk→embed→retrieve→rerank→generate con grounding; modelo/metric justificados; citaciones y ruta de refusal. [EXPLICIT]
  • Límites: retrieval + generación aumentada; no fine-tuning del modelo base ni diseño de prompts genérico. [EXPLICIT]
  • Casos borde: corpus multilingüe → modelo de embedding multilingüe; retrieval vacío → refusal, no inventar. [DOC]
  • Supuestos: corpus disponible y un vector store/índice accesible. [SUPUESTO]
  • Trade-off: hybrid + rerank suben calidad a cambio de latencia/costo por consulta. [EXPLICIT]

Related Skills

  • ai-software-architecture — the system that embeds the RAG component. [EXPLICIT]
  • data-strategy — sourcing/governing the corpus. [EXPLICIT]
  • notebooklm-research, web-research — gather sources to ingest. [EXPLICIT]

Packet

Capas del packet, cargables bajo demanda (disciplina ICM: una capa por vez, nunca todas juntas): references/ guías de profundidad (cargar UNA por etapa) · knowledge/ cuerpo de conocimiento · prompts/ prompts listos · examples/ salida de ejemplo · agents/ subagentes del packet · assets/ recursos estáticos.

Repository
JaviMontano/claude-plugins
Last updated
First committed

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.