CtrlK
BlogDocsLog inGet started
Tessl Logo

custom-index-eval

Evaluates Fusion MCP search quality against eval/index domain files. USE FOR: eval core, eval all, validate MCP index recall, check search freshness, verify documented framework patterns. DO NOT USE FOR: authoring eval patterns, editing docs, CI automation, or application-development search answers.

69

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

82%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, actionable skill body with a clean workflow, explicit edge-case handling, and a verified report-template reference. The main defect is the missing `agents/query-judge.md` file referenced in step 3, which is a dangling reference and weakens both actionability and progressive disclosure.

Suggestions

Add the referenced `agents/query-judge.md` to the bundle (or remove the reference and inline the judging criteria), since step 3 tells Claude to prefer a sub-agent whose file does not exist.

Include a short worked example of one domain-file query with its `must`/`should` expectations and the corresponding verdict, so the parse-then-judge contract is unambiguous.

Add a brief fallback rule for malformed domain files or partial MCP failures during `eval all` (e.g., report the domain as skipped vs. failed) to close the last workflow gap.

DimensionReasoningScore

Conciseness

The body is lean and imperative throughout ("Skip empty files or files without `##` query headings", "Verdicts are `pass`, `partial`, or `fail`; never inflate ambiguous results") with no explanation of concepts Claude already knows and every line earning its place. The trigger-phrase list adds terms beyond the frontmatter rather than duplicating it.

5 / 5

Actionability

Concrete, executable rules are given: file mapping (`eval/index/<domain>.md`), heading-parse conventions (`- must ...` / `- should ...` bullets), named search tools (`mcp_fusion_search_framework`, then `mcp_fusion_search`), verdict vocabulary, and a real report template. The gap is the reference to an `agents/query-judge.md` sub-agent that is not present in the bundle, though the inline fallback keeps the flow recoverable.

4 / 5

Workflow Clarity

The four-step sequence (resolve → parse → judge → report) is clearly ordered with edge-case handling (skip empty files, sub-agent fallback, "Stop clearly if MCP is unavailable or rate-limited"), which function as checkpoints. Minor gaps: no guidance for malformed domain files or partial failures midway through an `eval all` run, and aggregation across domains is only implied by the report instruction.

4 / 5

Progressive Disclosure

The ~58-line body is well-sectioned (When To Use, Inputs, Workflow, Safety, Expected Output) and points to `assets/report-template.md`, a real one-level-deep bundle file that exists. However, `agents/query-judge.md` is referenced but absent from the bundle, a dangling navigation path.

4 / 5

Total

17

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concrete, with explicit USE FOR trigger phrases and a DO NOT USE FOR exclusion list that bounds the niche. The only minor weakness is that trigger phrasing crowds out a fuller statement of what the skill's output looks like (verdicts, pass-rate reports).

DimensionReasoningScore

Specificity

Names the domain ("Fusion MCP search quality against eval/index domain files") and lists several concrete actions — "validate MCP index recall", "check search freshness", "verify documented framework patterns" — but these are framed as trigger phrases rather than capability statements, leaving minor gaps in what the skill actually produces (verdicts, pass rates).

4 / 5

Completeness

Explicitly answers both: what ("Evaluates Fusion MCP search quality against eval/index domain files") and when ("USE FOR: eval core, eval all, validate MCP index recall...") with concrete trigger phrases, plus an explicit "DO NOT USE FOR" exclusion list.

5 / 5

Trigger Term Quality

"eval core", "eval all", "validate MCP index recall", "check search freshness" are natural commands a user of this tooling would say. A few natural variations the body itself lists ("check MCP index accuracy", "is the index stale?") are missing from the description, so coverage is good but not comprehensive.

4 / 5

Distinctiveness Conflict Risk

Clear niche (evaluating a specific Fusion MCP search index) with distinct triggers; the "DO NOT USE FOR: authoring eval patterns, editing docs, CI automation, or application-development search answers" boundary sharply reduces overlap with adjacent skills.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_field

'metadata' should map string keys to string values

Warning

Total

15

/

16

Passed

Repository
equinor/fusion-framework
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.