CtrlK
BlogDocsLog inGet started
Tessl Logo

he-improve

Improve existing Harness Engineering implementations or workflows with evidence-backed changes. Use when users ask for targeted enhancement of shipped or drafted work.

55

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Failed to scan

The risk profile of this skill

Fix and improve this skill with Tessl

tessl review fix ./Plugins/harness-engineering/fixtures/budget-archive/2026-04-21/deferred-store/skills/team_automation/he-improve/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-structured, reasonably lean process skill with clear sequencing and validation, but its guidance stays directional rather than executable and its progressive-disclosure claims are undermined by missing/broken bundle references. Strongest on workflow clarity; weakest on actionability and reference integrity.

Suggestions

Add concrete specifics to the procedure: an example optimization-spec schema, the actual command that runs the measurement harness, and the on-disk path/format for experiment state, to lift actionability above directional hints.

Fix or remove the broken references: either ship the missing ./assets/ icons and references/session-evidence-contract.md (at a sane relative path) or delete the 'Full Context' and step-2 links so the entrypoint does not promise material that is absent.

Inline one or two validation checkpoints directly into the Procedure steps (e.g., a 'Validate spec before experimentation' gate at step 1) instead of relying on a separate Validation section, to tighten the feedback loop.

DimensionReasoningScore

Conciseness

The body is mostly lean bullet lists with no padding explaining concepts Claude already knows, fitting the score-4 anchor of 'efficient; minor instances of over-explanation that could be trimmed.' It is not score 5 because the baseline-metrics theme recurs across Philosophy, Validation, Anti-patterns, and Constraints, creating minor redundancy that could be tightened.

4 / 5

Actionability

The procedure gives a clear named-artifact sequence ('optimization spec', 'measurement harness', 'collector output') but no concrete commands, spec format, file paths, or executable examples, matching the score-3 anchor of 'some concrete guidance but incomplete; missing key details.' It is not score 4 because steps like 'Run bounded iterations with explicit measurement gates' remain directional rather than executable.

3 / 5

Workflow Clarity

The 8-step Procedure is clearly sequenced and paired with an explicit Validation section, a fail-fast gate, and a verify-write checkpoint (step 7), matching the score-4 anchor of 'clear sequence with most checkpoints present.' It is not score 5 because validation lives in a separate section rather than inline per-step, and the keep/revise/discard feedback loop is stated but not tied to a re-run step.

4 / 5

Progressive Disclosure

The body is well-organized into clear sections and signals references explicitly, but the referenced bundle files do not exist: './assets/icon-small.png', './assets/icon-large.png', and '../../../../../../references/session-evidence-contract.md' all resolve to nothing, and the latter is a deeply nested 6-level-up path. This fits the score-3 anchor of 'some structure; references present but not clearly signaled / could be better organized.' It is not score 4 because the claimed archived-reference layer the entrypoint depends on is effectively absent.

3 / 5

Total

14

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description cleanly answers both what and when with a clear niche, but its action vocabulary is narrow and it lists only one generic action rather than multiple concrete capabilities. It is solid but not comprehensive.

Suggestions

Enumerate 2-3 concrete actions (e.g., 'tune retry workflows, compare bounded alternatives, route proven improvements to the next stage') to raise specificity from generic 'improve'.

Add natural synonyms and surface forms users would say (e.g., 'tune', 'optimize', 'shipped workflow', 'drafted plugin') to broaden trigger-term coverage.

Add one more 'Use when...' trigger clause covering the session-evidence path so the when-guidance matches the body's actual scope.

DimensionReasoningScore

Specificity

The description names the domain ('Harness Engineering implementations or workflows') and one concrete action ('Improve ... with evidence-backed changes'), but does not enumerate multiple specific actions, matching the score-3 anchor of domain plus 1-2 actions that are not comprehensive. It is above score 2 because 'evidence-backed changes' and 'targeted enhancement' are more concrete than a bare generic verb, but below score 4 because no list of distinct capabilities is provided.

3 / 5

Completeness

It explicitly answers both 'what' ('Improve existing Harness Engineering implementations or workflows with evidence-backed changes') and 'when' ('Use when users ask for targeted enhancement of shipped or drafted work'), matching the score-4 anchor of both present with the 'when' reasonably explicit. It is not score 5 because the trigger clause offers a single phrasing rather than the comprehensive, multi-phrase trigger coverage of the score-5 example.

4 / 5

Trigger Term Quality

Natural trigger phrases like 'targeted enhancement of shipped or drafted work' and 'improve' are present, but the vocabulary is narrow with no synonyms or variations, fitting the score-3 anchor of 'some relevant keywords but missing common variations.' It is not score 4 because coverage lacks the breadth of natural terms a user might actually say.

3 / 5

Distinctiveness Conflict Risk

The 'Harness Engineering' niche term and 'shipped or drafted work' scoping make it mostly distinct with only minor overlap risk against general optimization skills, matching the score-4 anchor. It is not score 5 because the verb 'improve/enhance' is generic enough that overlap with other improvement-tuning skills remains possible.

4 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

relative_links

Relative link issues: 2 missing, 1 suspicious

Warning

Total

14

/

16

Passed

Repository
jscraik/Agent-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.