CtrlK
BlogDocsLog inGet started
Tessl Logo

testland/ragas-evaluation

Authors and runs Ragas - RAG-pipeline evaluation framework with metrics organized into RAG (Faithfulness, Response Relevancy, Context Precision/Recall, Context Entities Recall, Noise Sensitivity), Natural Language Comparison (Factual Correctness, Semantic Similarity, BLEU/ROUGE/CHRF/Exact Match), Agents/Tool-Use (Topic Adherence, Tool Call Accuracy/F1, Agent Goal Accuracy), General Purpose (Aspect Critic, Rubrics-based Scoring), Nvidia (Answer Accuracy, Context Relevance, Response Groundedness), and Summarization. Use when the user evaluates a RAG pipeline (retriever + generator) and needs the deepest metric variety in the OSS LLM-eval space.

67

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Overview
Quality
Evals
Security
Files

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, efficient body: install, a runnable custom-metric quickstart, a metric-selection step delegated to a real references file, dataset schema table, and CI plus anti-pattern guidance. The gaps are the doc-deferred integration step (no concrete LangChain/LlamaIndex wiring code) and the absence of an explicit validation/smoke-test checkpoint in the workflow.

Suggestions

Add a minimal, runnable LangChain or LlamaIndex integration snippet to Step 5 (even a 5-line testset-generation example) instead of deferring entirely to docs.ragas.io.

Add a Step 2.5 or post-install smoke check (e.g., run one metric on a single-row dataset and confirm a numeric score before scaling up) to give the workflow an explicit validation checkpoint.

Make the CI example self-contained by showing how `dataset` is constructed (one-line testset or EvaluationDataset.from_dict) so the snippet is copy-paste runnable.

DimensionReasoningScore

Conciseness

The body is lean and assumes competence (no explanation of what RAG or evaluation is); the "Pick 3 - 5 per pipeline" guidance appears in both Step 3 and the anti-patterns table, and "Per [rg-gh]" is cited repeatedly, so minor trims are possible. Not 5 due to those small redundancies; not 2-3 because padding is the exception rather than the rule.

4 / 5

Actionability

Install commands and the verbatim DiscreteMetric quickstart are copy-paste ready, and the CI snippet gives a concrete pattern. Not 5 because Step 5 defers integration entirely to external docs ("Consult the per-framework integration docs... when wiring") with no code, and the CI example references an undefined `dataset`; not 3 because most steps provide executable guidance.

4 / 5

Workflow Clarity

Clear Steps 1-6 sequence with a dataset-shape table, anti-pattern/fix table, and CI threshold assertions as checkpoints. Not 5 because there is no explicit validation or smoke-run step after install (e.g., verify the metric runs on a sample row) or feedback loop for a failed eval; not 3 because the sequence and most checkpoints are present and unambiguous.

4 / 5

Progressive Disclosure

SKILL.md is a concise overview that delegates the full metric catalog to references/metrics.md (verified to exist), clearly signaled and one level deep: "Full per-family catalog with each metric's use: [references/metrics.md](references/metrics.md)". Navigation is easy with no inlined bulk content.

5 / 5

Total

17

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, explicit description: it names the framework and its actions in third person, comprehensively catalogs the metric families, and includes a clear "Use when..." clause tied to RAG pipeline evaluation. The main weakness is verbosity from enumerating every metric family, which slightly widens the trigger surface relative to sibling eval skills.

DimensionReasoningScore

Specificity

"Authors and runs Ragas - RAG-pipeline evaluation framework" names the domain and concrete actions, with a comprehensive enumeration of specific metrics ("Faithfulness, Response Relevancy, Context Precision/Recall", "Tool Call Accuracy/F1", "Aspect Critic") in third-person voice. Not 5 because it enumerates a metric catalog rather than multiple distinct actions; not 3 because coverage goes well beyond 1-2 actions.

4 / 5

Completeness

Explicitly answers what ("Authors and runs Ragas - RAG-pipeline evaluation framework with metrics organized into...") and when ("Use when the user evaluates a RAG pipeline (retriever + generator) and needs the deepest metric variety in the OSS LLM-eval space") with concrete trigger phrases.

5 / 5

Trigger Term Quality

Includes natural phrases users would say: "evaluates a RAG pipeline (retriever + generator)", "Ragas", "LLM-eval", plus concrete metric names (Faithfulness, Context Recall). Not 5 because common variations like "RAG evaluation" or "eval my retriever" phrasings are absent; not 3 because domain terms and synonyms are largely covered.

4 / 5

Distinctiveness Conflict Risk

Clear niche (RAG-pipeline evaluation via Ragas) with distinct trigger phrasing, but the breadth of metric families beyond RAG ("Agents/Tool-Use", "Summarization", "General Purpose") creates minor overlap risk with general-purpose LLM-eval skills. Not 5 due to that overlap; not 3 because the trigger is firmly scoped to RAG pipelines.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Reviewed

Table of Contents