Configures and runs LLM evaluation using Promptfoo framework. Use when setting up prompt testing, creating evaluation configs (promptfooconfig.yaml), writing Python custom assertions, implementing llm-rubric for LLM-as-judge, or managing few-shot examples in prompts. Triggers on keywords like "promptfoo", "eval", "LLM evaluation", "prompt testing", or "model comparison".
87
81%
Does it follow best practices?
Impact
97%
1.59xAverage score across 3 eval scenarios
Passed
No findings from the security scan
| Run | Type | Date | Status |
|---|---|---|---|
baseline vs usage-spec With / without contextCompleted | With / without context | Completed |
019cbea9-39b9-717b-be92-b0d66e668bc4
Run
The run is available. Its result stats will appear here when they are ready.
bb6ad55
Table of Contents
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.