CtrlK
BlogDocsLog inGet started
Tessl Logo

he-improve

Improve existing Harness Engineering skills, references, contracts, and evals from concrete evidence such as failed evals, repeated review findings, usage traces, or documented regressions. Use when a bounded hardening pass is required; do not use for speculative redesign.

66

Quality

78%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

Fix and improve this skill with Tessl

tessl review fix ./Plugins/harness-engineering/skills/he-improve/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

72%Weight 40%Scale 1-3

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a dense, actionable instruction set with strong progressive disclosure pointing to real local bundle files. Its weaknesses are mechanical: truncated/dangling sentences in several sections and a duplicated step number that break conciseness and workflow sequence.

Suggestions

Repair the truncated sentences at the end of Philosophy (line 16), When Not to Use (line 28), Procedure (line 61), and Validation (line 71) so each section ends with a complete thought.

Renumber the Procedure so the two steps both labeled 4 (lines 57 and 59) become 4 and 5, restoring an unambiguous sequence.

Convert the dangling reference-link lines (e.g., "- See references/hot-path-folded-context.md for folded philosophy detail.") into complete sentences or move them into a dedicated References subsection per section.

DimensionReasoningScore

Conciseness

Sections are terse and assume Claude's competence with no concept padding, but several lines are truncated mid-sentence (e.g., "Higher-priority instructions, command boundaries, and" at line 16, and dangling clauses at lines 28, 61, 71) and Procedure has two steps both numbered 4, so it could be tightened and repaired.

2 / 3

Actionability

It gives concrete, specific guidance for an instruction-only skill: a structured output schema (schema_version: 1 with named fields), an explicit side-effect taxonomy, git-staging contract, named validation gates, and fixed output sections (Routing, Evidence, Gaps, Patch, Validation, etc.).

3 / 3

Workflow Clarity

The procedure is sequenced and validation has an explicit fail-fast feedback loop ("stop at the first failed gate, fix or block it, then rerun") with before/after comparison, but the duplicated step 4 and truncated procedure/validation tails leave sequence gaps and implicit checkpoints.

2 / 3

Progressive Disclosure

The body is a lean overview that signals one-level-deep references to real local bundle files (references/hot-path-folded-context.md, references/contract.yaml, references/evals.yaml, assets/), with content appropriately split rather than inlined.

3 / 3

Total

10

/

12

Passed

Description

85%Weight 40%Scale 1-3

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, complete, and well-scoped with an explicit Use-when trigger and a clear negative boundary. Its main weakness is trigger-term naturalness: the leading trigger phrase is internal jargon rather than language a user would naturally say.

Suggestions

Rewrite the Use-when trigger in natural user phrasing (e.g., 'Use when a skill, reference, contract, or eval suite has failed validation or repeated review findings and needs a bounded improvement') so users would actually say it.

Add common trigger variations such as 'hardening pass', 'failed eval', 'review finding', or 'regression' as explicit keywords to improve trigger-term coverage.

DimensionReasoningScore

Specificity

Names a concrete domain ("Harness Engineering skills, references, contracts, and evals") and multiple specific motivating evidence types ("failed evals, repeated review findings, usage traces, or documented regressions"), matching the score-3 anchor of listing multiple specific concrete actions.

3 / 3

Completeness

It explicitly answers both what ("Improve existing Harness Engineering skills, references, contracts, and evals from concrete evidence") and when ("Use when a bounded hardening pass is required") with an explicit Use-when clause.

3 / 3

Trigger Term Quality

It includes relevant terms (failed evals, review findings, regressions) but the primary trigger "a bounded hardening pass is required" leans on internal jargon rather than natural user phrasing, and common variations are missing.

2 / 3

Distinctiveness Conflict Risk

The niche is specific (bounded evidence-driven hardening of HE surfaces) and the negative carve-out "do not use for speculative redesign" plus handoff routing to he-fix-bugs keeps it distinct from adjacent skills.

3 / 3

Total

11

/

12

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
jscraik/Agent-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.