CtrlK
BlogDocsLog inGet started
Tessl Logo

benchmark-workflow

Run, diagnose, or change Xberg extraction benchmarks, quality scoring, benchmark fixtures, artifact contracts, and independently sourced ground truth. Load for the Benchmarks workflow or benchmark-harness work, not ordinary unit tests.

72

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A tight, well-organized instruction-only skill that assumes Claude's competence and provides concrete file paths, commands, and field names. Its main gap is the absence of worked command examples and a fully gated step sequence.

Suggestions

Add one or two concrete command invocations (e.g. the exact validate-gt call and how to run the affected adapter task) so guidance is copy-paste ready.

Convert the "Diagnosing runs" bullets into a short numbered procedure with an explicit validation gate before dispatching the remote workflow.

DimensionReasoningScore

Conciseness

Lean and efficient throughout: tight bullets, no explaining of concepts Claude already knows, and every line delivers actionable guidance, matching the every-token-earns-its-place anchor.

5 / 5

Actionability

Highly concrete pointers (specific paths like tools/benchmark-harness/, the validate-gt command, generate_markdown_gt.py, and enumerated ground_truth.source values), but lacks worked command invocations or examples showing exact usage for the common cases.

4 / 5

Workflow Clarity

Clear sequenced diagnostic priority (separate infra vs extraction failures; inspect per-adapter before aggregate; compare OCR pages before word counts) with validation checkpoints and a ground-truth feedback loop, but it is a principles/checklist guide rather than a fully gated numbered workflow.

4 / 5

Progressive Disclosure

Under 50 lines with no bundle files needed, organized into clear sections (Ground-truth integrity, Diagnosing runs) with concise inline pointers to repo paths, satisfying the simple-skill exception for a well-organized overview.

5 / 5

Total

18

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, third-person description that crisply states capabilities and explicit load conditions with a useful negative boundary. It is comprehensive on what/when and distinctiveness, with only minor room to add synonyms and file extensions for trigger terms.

Suggestions

Add a file-extension or path cue (e.g. ".github/workflows/benchmarks.yaml") so users referencing the workflow file by name trigger the skill.

Include a couple of natural synonyms such as "benchmark scores" or "benchmark regression" to broaden trigger coverage.

DimensionReasoningScore

Specificity

Lists multiple concrete actions ("Run, diagnose, or change") applied to an enumerated, comprehensive set of domain objects (benchmarks, quality scoring, fixtures, artifact contracts, ground truth), matching the comprehensive-coverage anchor.

5 / 5

Completeness

Explicitly answers both what ("Run, diagnose, or change Xberg extraction benchmarks...") and when ("Load for the Benchmarks workflow or benchmark-harness work, not ordinary unit tests") with concrete trigger phrases and a negative boundary.

5 / 5

Trigger Term Quality

Good natural keyword coverage ("Benchmarks workflow", "benchmark-harness work", "not ordinary unit tests") that users in this repo would say, but lacks synonyms and file extensions (e.g. .yaml) that would push it to comprehensive.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (Xberg extraction benchmarks / benchmark-harness) with distinct triggers and an explicit exclusion ("not ordinary unit tests"), giving minimal conflict risk with other skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
xberg-io/xberg
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.