CtrlK
BlogDocsLog inGet started
Tessl Logo

promptfoo-evaluation

Configures and runs LLM evaluation using Promptfoo framework. Use when setting up prompt testing, creating evaluation configs (promptfooconfig.yaml), writing Python custom assertions, implementing llm-rubric for LLM-as-judge, or managing few-shot examples in prompts. Triggers on keywords like "promptfoo", "eval", "LLM evaluation", "prompt testing", or "model comparison".

87

1.59x
Quality

81%

Does it follow best practices?

Impact

97%

1.59x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

75%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A highly actionable body: all guidance is executable code and config with a sensible setup-to-run flow, a preview-first validation step, and a troubleshooting section. The main costs are token-weight from triple-stated gotchas and body/reference duplication, plus one dangling reference (./tiaogaoren/) that breaks navigation at the point it promises a full worked example.

Suggestions

De-duplicate the relay/apiBaseUrl and maxConcurrency/commandLineOptions gotchas: state each once in its primary section and keep only one-line pointers in Troubleshooting (or move the pointers into the troubleshooting entries) to lift conciseness.

Remove or fix the dangling "See: ./tiaogaoren/ (example project root)" reference — either include the directory in the bundle or replace it with the inline structure snippet already shown, so progressive disclosure navigation is fully intact.

Replace the duplicated Echo Provider section with a two-line pointer to references/promptfoo_api.md, and add an explicit line telling the reader that the bundled scripts/metrics.py implements the assertion helpers shown inline.

DimensionReasoningScore

Conciseness

The body is mostly efficient, executable examples, but the relay/apiBaseUrl gotcha is stated three times (in the llm-rubric section, its best-practices list, and Troubleshooting) and the maxConcurrency/commandLineOptions rule likewise three times, and the Echo Provider section duplicates content already in references/promptfoo_api.md — fitting the score-3 anchor ("could be tightened") rather than score 4's "minor instances".

3 / 5

Actionability

Every section provides copy-paste-ready YAML, Python, or bash (full promptfooconfig.yaml, working get_assert/custom_check/strip_tags functions, echo-provider preview config, CLI commands with flags), fully covering the common cases per the score-5 anchor.

5 / 5

Workflow Clarity

The init → config → prompts/tests → assertions → preview → run → view sequence is logically ordered with an explicit validation checkpoint ("Use echo provider first to verify structure") and a troubleshooting section for error recovery, but validation is recommended per-topic rather than woven into one explicit numbered workflow, leaving the minor gaps of the score-4 anchor; it is not score 5 because there is no single validate-then-proceed checklist.

4 / 5

Progressive Disclosure

Headers are clear, the reference (references/promptfoo_api.md) is real and one level deep with a clearly signaled link, but the body's "See: ./tiaogaoren/" points at a directory that does not exist in the bundle, scripts/metrics.py is never explicitly offered as a bundled asset, and the Echo Provider content is duplicated between body and reference — minor organization gaps per the score-4 anchor rather than score 5's clean split.

4 / 5

Total

16

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that names the framework and enumerates concrete actions with explicit what/when guidance and a natural trigger list. Its only weaknesses are a trigger list missing common synonyms and a couple of generic keywords ("eval", "model comparison") that slightly raise conflict risk with adjacent evaluation tooling skills.

DimensionReasoningScore

Specificity

The description lists five concrete, named actions — "Configures and runs LLM evaluation using Promptfoo framework", "creating evaluation configs (promptfooconfig.yaml)", "writing Python custom assertions", "implementing llm-rubric for LLM-as-judge", "managing few-shot examples" — covering the domain comprehensively with a specific artifact name, matching the score-5 anchor rather than the score-4 anchor ("minor gaps in coverage").

5 / 5

Completeness

Both questions are explicitly answered: the "what" in the opening sentence and the "when" via "Use when setting up prompt testing..." plus an explicit trigger-keyword list, with concrete trigger phrases — the exact shape of the score-5 anchor; it is not score 4 because the "when" clause needs no added explicitness.

5 / 5

Trigger Term Quality

Triggers "promptfoo", "eval", "LLM evaluation", "prompt testing", "model comparison" are natural phrases users would say, but the list misses common variants (e.g., "benchmark", "assertion", "grading", "promptfooconfig"), so it fits the score-4 anchor ("a few natural terms missing") rather than score 5's synonym-and-extension breadth.

4 / 5

Distinctiveness Conflict Risk

The named framework (Promptfoo) gives a clear niche, but the generic triggers "eval" and "model comparison" create minor overlap risk with other LLM-evaluation or model-benchmarking skills, fitting the score-4 anchor ("minor overlap risk with closely related skills") rather than score 5's "minimal conflict risk".

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
daymade/claude-code-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.