CtrlK
BlogDocsLog inGet started
Tessl Logo

test-boost-module

Live-test a Harbor Boost module by sending a real prompt through llamacpp via pi and validating the output. Use when asked to test a boost module, verify a module works, check module behavior, QA a boost module, or confirm a module's effect on LLM output.

72

Quality

88%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

High

Do not use without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An excellent, execution-ready skill body: fully concrete commands, a control-run experimental design, a validation checklist, and thorough troubleshooting feedback loops, all in a lean single file. The only improvement available is trimming the near-duplicate harbor launch command listings.

DimensionReasoningScore

Conciseness

The body is dense and assumes competence — no tutorials on what llamacpp or docker are — with every section carrying operational value (the module-type prompt table, argument-order note, troubleshooting). Minor trimmable redundancy: the full 'harbor launch ... pi -p --no-tools --no-session' command is repeated nearly verbatim in 'The Command' and again twice in 'Running the Test' steps 2-3. Anchor 4 (efficient with minor instances that could be trimmed) fits better than 5.

4 / 5

Actionability

Every instruction is copy-paste executable: the 'harbor launch --workflow <module> --model ... pi -p --no-tools --no-session' command, the curl+python3 model-listing one-liner with the auth header and port derivation, 'docker logs harbor.boost --tail 20', and named default models. Placeholders are legitimate parameterization, not pseudocode, and common cases (unknown module type, indistinguishable output) are covered.

5 / 5

Workflow Clarity

The sequence is explicit and ends in validation: prerequisites with health-check gating, model selection, test run, control run on the same model/prompt, output comparison, then a five-item PASS/FAIL checklist. Feedback loops are present throughout — pick a better prompt, check Boost logs, restart Boost and retry, timeout/smaller-model fallback — matching the anchor 5 pattern of explicit validation plus error-recovery loops.

5 / 5

Progressive Disclosure

No bundle directories (references/, scripts/, assets/) exist and the body references no external files, so there is nothing mis-split or nested; scored against the actual single-file structure. At ~130 lines with well-organized sections (Prerequisites, Command, Prompt choice, Running, Checklist, Troubleshooting), everything is appropriately inlined for SKILL.md and navigation is trivial, satisfying the no-external-references exception.

5 / 5

Total

19

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete third-person capability statement paired with an explicit multi-phrase 'Use when' trigger list. Its only weaknesses are slight genericness in a couple of trigger phrases and a missing hint at the control-run comparison method.

DimensionReasoningScore

Specificity

Phrases like 'Live-test a Harbor Boost module by sending a real prompt through llamacpp via pi and validating the output' name several concrete actions (live-test, send a real prompt, validate output) in third person, with only a minor gap — it never hints at the with/without-module comparison that defines the method. Not 3 because more than 1-2 actions are specifically named; not 5 because coverage is not comprehensive.

4 / 5

Completeness

Both halves are explicit: 'what' is 'Live-test a Harbor Boost module by sending a real prompt through llamacpp via pi and validating the output' and 'when' is the concrete 'Use when asked to test a boost module, verify a module works, ...' trigger list. This matches the anchor 5 example's structure exactly; nothing is merely implied.

5 / 5

Trigger Term Quality

'test a boost module, verify a module works, check module behavior, QA a boost module, or confirm a module's effect' gives good natural-phrase coverage (test/verify/check/QA/confirm). Not 5 because a few natural synonyms (e.g. 'smoke test', 'run/integration test a module', file-extension-style identifiers) are missing.

4 / 5

Distinctiveness Conflict Risk

The 'what' anchors a clear niche (Harbor Boost, llamacpp, pi), but trigger phrases like 'verify a module works' and 'check module behavior' are generic enough to fire for non-Boost module verification. Mostly distinct with minor overlap risk, matching anchor 4 rather than 5.

4 / 5

Total

17

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
av/harbor
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.