CtrlK
BlogDocsLog inGet started
Tessl Logo

phoenix-evals

Build and run evaluators for AI/LLM applications using Phoenix.

58

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./.agents/skills/phoenix-evals/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an excellent, lean navigation hub with strong progressive disclosure and clear multi-step workflows. Its main weakness is actionability: the body itself contains no executable code or commands, relying entirely on referenced files for concrete steps.

Suggestions

Add one tiny inline code snippet or concrete command in the Quick Reference (e.g., a minimal evaluator definition or the experiment run call) so the body alone is partially executable.

Make validation checkpoints explicit within at least the Building Evaluator and Gating CI workflows (e.g., a "validate >80% TPR/TNR before gating" step) to push workflow clarity to 5.

Annotate which workflow entry point to start from based on user intent (new vs. existing project) to further aid navigation.

DimensionReasoningScore

Conciseness

The body is a lean hub of tables and arrow-sequenced workflow pointers with no concept-explaining padding; it assumes Claude's competence and every token earns its place, matching anchor 5.

5 / 5

Actionability

It gives concrete, task-to-file guidance (Quick Reference table, named workflow files) but no executable code or commands, so execution is incomplete; anchor 3 fits better than 4 because there are no copy-paste-ready steps in the body itself.

3 / 5

Workflow Clarity

Five named workflows use arrow-sequenced steps (e.g., "observe-tracing-setup → error-analysis → axial-coding → evaluators-overview") giving a clear sequence, matching anchor 4; it does not reach 5 because validation checkpoints/feedback loops are only implied via the Key Principles table rather than explicit per workflow.

4 / 5

Progressive Disclosure

It is a clear overview with well-signaled one-level-deep references (task table, category table, workflows) pointing to 37 real reference files verified on disk; content is appropriately split and easy to navigate, matching anchor 5.

5 / 5

Total

17

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names a clear, distinct niche but provides only minimal actions and no "when to use" trigger guidance, which caps completeness. Adding a Use when... clause and more natural trigger synonyms would raise it.

Suggestions

Append a "Use when..." clause naming concrete trigger scenarios, e.g. "Use when evaluating LLM outputs, building LLM-as-a-judge evaluators, or running AI experiments in Phoenix."

Expand the action list beyond build/run (e.g., validate evaluators against human labels, gate CI with evals, sample traces for review) to improve specificity.

Add natural synonym/extension terms users say ("LLM evals", "LLM-as-a-judge", "AI observability", "eval datasets") for better trigger coverage.

DimensionReasoningScore

Specificity

"Build and run evaluators for AI/LLM applications using Phoenix" names the domain and two concrete actions (build, run) but is not comprehensive, matching anchor 3; it does not reach 4 because only two minimal actions are listed, and stays above 2 because the actions are concrete rather than generic.

3 / 5

Completeness

It gives a clear "what" (build and run evaluators for AI/LLM apps using Phoenix) but has no "Use when..." clause, so per the missing-trigger-guidance cap it stays at anchor 3 and cannot reach 4 or 5.

3 / 5

Trigger Term Quality

It includes relevant terms ("evaluators", "AI/LLM applications", "Phoenix") that a user might say, but misses common variations like "LLM evals", "LLM-as-a-judge", or "AI observability", fitting anchor 3 rather than 4's broader keyword coverage.

3 / 5

Distinctiveness Conflict Risk

"Evaluators for AI/LLM applications using Phoenix" is a fairly distinct niche with only minor overlap risk against general observability skills, matching anchor 4; it does not reach 5 because it is very short and lacks the explicit trigger phrases of the anchor-5 example.

4 / 5

Total

13

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
Arize-ai/phoenix
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.