CtrlK
BlogDocsLog inGet started
Tessl Logo

harness-eval

This skill should be used when the user asks to "test the harness", "run integration tests", "validate features with real API", "test with real model calls", "run agent loop tests", "verify end-to-end", or needs to verify OpenHarness features on a real codebase with actual LLM calls.

67

Quality

84%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is harness-eval in HKUDS/OpenHarness

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable eval-workflow skill with an exemplary sequence, validation checkpoints, and clean one-level-deep reference structure. Its main weakness is redundancy: the max_turns and timeout-triage guidance recurs across four sections and could be consolidated into a single rule.

Suggestions

State the max_turns rule once (e.g., in section 5) and delete its repetitions in section 2, the pitfalls list, and section 7, keeping only a one-line cross-reference where needed.

Merge the timeout/failure-triage guidance that currently appears in both section 7 and Common Pitfalls into the section 7 failure-classification list.

Complete the section 3 smoke-check snippet (the "# Point config loader at this file..." line) or point to the exact template in references/test-patterns.md so it is copy-paste runnable.

DimensionReasoningScore

Conciseness

The body has no filler or explanations of concepts Claude already knows, but the max_turns guidance is repeated four times ("do not artificially lower max_turns", "Keep max_turns=200", the pitfall entry, "First check whether max_turns was manually set too low") and timeout/failure-triage guidance is duplicated across sections 5, 7, and Common Pitfalls, so it could be meaningfully tightened.

3 / 5

Actionability

Concrete and mostly executable throughout (clone/export/install commands, a smoke-check snippet, explicit test-run commands, an interpretation table), but the section 3 snippet ends in a comment placeholder ("# Point config loader at this file, then run BashTool...") and the core make_engine/collect helpers are only sketched inline, deferring to the reference file.

4 / 5

Workflow Clarity

A numbered 7-step sequence carries explicit validation checkpoints (sandbox smoke check through the real adapter path, `which srt/bwrap/rg`, "inspect tool call lists and output files, not just model text"), a failure-classification feedback loop ("classify it before changing code"), per-scenario success criteria, and a feature coverage checklist.

5 / 5

Progressive Disclosure

The body defers code templates to references/test-patterns.md (signaled inline in section 4) and lists both real, one-level-deep, non-nested reference files (test-patterns.md, feature-matrix.md) with descriptions in a dedicated Additional Resources section; the split of detail between SKILL.md and the references is appropriate.

5 / 5

Total

17

/

20

Passed

Description

86%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A well-constructed description with an explicit use-when clause, six natural trigger phrasings, and a clearly stated purpose. Its only notable gap is that the capability statement is a single high-level action rather than a concrete enumeration of what the skill actually does.

Suggestions

Enumerate the core actions in the what-clause (e.g., "run real multi-turn agent loops against a cloned unfamiliar repo and verify tool execution") to raise specificity.

Qualify the generic triggers (e.g., "run integration tests against OpenHarness") so they cannot fire for ordinary testing requests unrelated to this product.

DimensionReasoningScore

Specificity

The description names the domain and essentially one concrete action ("verify OpenHarness features on a real codebase with actual LLM calls") rather than enumerating several specific actions such as cloning an unfamiliar repo, running multi-turn agent loops, and inspecting tool traces, so it matches "domain plus 1-2 concrete actions" rather than the several-action level above.

3 / 5

Completeness

It explicitly answers both questions: an explicit "This skill should be used when the user asks to..." clause with concrete trigger phrases (when), and a clear statement of what it does — verify OpenHarness features on a real codebase with actual LLM calls (what).

5 / 5

Trigger Term Quality

It provides comprehensive natural trigger phrasings with synonym coverage: "test the harness", "run integration tests", "validate features with real API", "test with real model calls", "run agent loop tests", "verify end-to-end" — a user needing this skill would naturally say one of these.

5 / 5

Distinctiveness Conflict Risk

The niche is clear and product-specific (OpenHarness + real LLM calls + real codebase), but generic triggers like "run integration tests" or "verify end-to-end" could plausibly fire in non-OpenHarness testing contexts, giving minor overlap risk with a general test-running skill.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
ProwlrBot/prowlr-cli
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.