CtrlK
BlogDocsLog inGet started
Tessl Logo

config-evals

Builds and maintains configuration-based evaluations on a workflow with the eval-config tool. Use when the user asks to set up, add, view, change, or remove an evaluation, score, grade, or judge a workflow's output, or measure answer quality against a test dataset. This is the only eval form Instance AI handles — it does not touch on-canvas evaluation nodes.

69

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, lean instruction skill: concrete tool-call guidance, a clear guarded procedure, and a properly disclosed single reference file for detail. Weaknesses are minor — some within-body repetition of the trigger and expression-form rules, and no full worked example inline.

DimensionReasoningScore

Conciseness

The body is lean and imperative with no background-concept padding, but a few points are stated twice: the 'never a trigger' start-node rule appears in both 'What a Config Eval Needs' and step 2, and the '={{ $json.<name> }}' expression form plus its literal-text failure mode appear in both 'Metrics' and the 'Expression fields must begin with =' subsection. Mostly efficient with minor duplication that could be trimmed fits the 4 anchor rather than the every-token-earns-its-place 5.

4 / 5

Actionability

Concrete, executable guidance throughout — specific tool calls ('data-tables(action="list")', 'eval-config' with action="create"/"update"), named config fields, and correct/wrong expression examples — but the body contains no complete copy-paste example of a full create call (those are deferred to the reference playbook). This matches 'mostly executable guidance with minor gaps' rather than the fully copy-paste-ready 5 anchor.

4 / 5

Workflow Clarity

The Default Procedure is a clear 6-step sequence with explicit guard points ('Never invent a dataTableId; use one returned by data-tables', 'call it and respect the result; do not ask for chat approval first', the start-node compile rule). It stops short of the 5 anchor because there are no error-recovery feedback loops (e.g., what to do when a run fails to compile or an approval is rejected), but the checkpoints that exist are explicit, keeping it above 3.

4 / 5

Progressive Disclosure

The ~105-line body is an appropriately scoped overview, and detailed recipes and worked examples are split into a single one-level-deep reference (references/config-eval-playbook.md), which exists in the bundle and is clearly signaled at the end with what it contains ('tool-call recipes, worked examples, and output shapes'). The playbook itself references no further files, so navigation is flat and easy — matching the top anchor.

5 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concrete, with an explicit 'Use when' clause and a clear boundary against on-canvas evaluation nodes. The only weakness is synonym coverage of natural trigger terms.

DimensionReasoningScore

Specificity

The description enumerates multiple concrete actions — 'set up, add, view, change, or remove an evaluation, score, grade, or judge a workflow's output, or measure answer quality against a test dataset' — covering the full CRUD plus metric-verb surface of the skill. Coverage is comprehensive rather than having only minor gaps, so it matches the top anchor.

5 / 5

Completeness

It explicitly answers 'what' in third person ('Builds and maintains configuration-based evaluations on a workflow with the eval-config tool') and 'when' with a concrete 'Use when the user asks to...' clause enumerating trigger cases. Both halves are explicit with concrete trigger phrases, matching the 5 anchor.

5 / 5

Trigger Term Quality

Natural user phrasings are well covered ('set up... an evaluation', 'score, grade, or judge a workflow's output', 'measure answer quality', 'test dataset'), but common synonyms such as 'eval', 'benchmark', 'assess', or 'test my workflow' are absent. Good keyword coverage with a few natural terms missing fits the 4 anchor, and it is clearly above 3 given the breadth that is present.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche (config-based evals on Instance AI workflows via eval-config) and adds an explicit boundary — 'the only eval form Instance AI handles — it does not touch on-canvas evaluation nodes' — minimizing conflict with adjacent skills. Distinct triggers plus a stated boundary fit the top anchor.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
n8n-io/n8n
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.