CtrlK
BlogDocsLog inGet started
Tessl Logo

gan-style-harness

GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously. Based on Anthropic's March 2026 harness design paper. Use when a feature should be built autonomously through generator and evaluator iteration until it clears a quality bar.

55

Quality

63%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/gan-style-harness/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a genuinely actionable harness design: concrete commands, a complete configuration surface, an explicit feedback loop with validation thresholds, and a strong anti-patterns section. Its weaknesses are padding (conceptual framing, results tables, time-sensitive model-stage content) and progressive-disclosure failures — the shell script and prompt templates it instructs the reader to use are missing from the bundle.

Suggestions

Create the referenced bundle files (scripts/gan-harness.sh, PLANNER_PROMPT.md, EVALUATOR_PROMPT.md) or remove/rewrite the usage sections that depend on them, since every invocation path currently points at nonexistent files.

Move the full evaluation rubric and the 'Evolution Across Model Capabilities' section into a references/ file, leaving SKILL.md as a lean overview with one-level-deep links.

Trim the 'Core Insight' quote and 'Results: What to Expect' table to one or two lines each, and quarantine dated model-version references ('Opus 4.5-class', 'March 2026') in a clearly-marked legacy section.

DimensionReasoningScore

Conciseness

The operational sections (usage, config, anti-patterns) are efficient, but the 'Core Insight' quote, the 'Evolution Across Model Capabilities' section, and the 'Results: What to Expect' table add context-heavy padding. Time-sensitive model references ('Opus 4.5-class', 'Opus 4.6-class', 'March 2026') are not quarantined in a deprecated/old-patterns section. Anchor 3 (mostly efficient, some unnecessary explanation) rather than 4's minor trimmable instances.

3 / 5

Actionability

Gives concrete, copy-paste-ready commands in all three usage modes (`/project:gan-build "..."`, `GAN_MAX_ITERATIONS=10 ./scripts/gan-harness.sh "..."`, and full `claude -p` prompts for the manual loop), plus a complete env-var and eval-mode table. Not 5 because the referenced `scripts/gan-harness.sh`, `PLANNER_PROMPT.md`, and `EVALUATOR_PROMPT.md` are absent from the bundle, leaving the primary invocation paths non-executable as shipped.

4 / 5

Workflow Clarity

The plan -> generate -> evaluate -> iterate loop is clearly sequenced ('Repeat steps 3-4 until pass threshold met') with explicit validation checkpoints: weighted scoring against the rubric, a configurable pass threshold, a max-iterations cap, and plateau detection ('stop and flag for human review'). Not 5 because project bootstrap, dev-server lifecycle (start/health-check/stop), and context-reset mechanics are only implied rather than sequenced as steps.

4 / 5

Progressive Disclosure

Sections are well-organized with clear headers, but the file is a monolithic ~280-line body: the full evaluation rubric and the model-evolution content sit inline where a reference file would serve, and navigation points to files that do not exist in the bundle (no references/, scripts/, or assets/ directories). Fits anchor 3 (some structure, content that should be separate is inline, references not backed by real files) rather than 4's well-placed split.

3 / 5

Total

14

/

20

Passed

Description

62%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description answers both what and when with a genuine explicit trigger clause, and its GAN/harness framing is fairly distinct. Weaknesses are jargon-heavy trigger phrasing, a thin list of concrete capabilities, and an abstract 'when' condition that may not match how users actually phrase the request.

Suggestions

Replace abstract trigger phrasing with natural user phrases, e.g. 'Use when the user asks to build an app or feature autonomously, keep iterating until quality is production-ready, or wants an adversarial generate-evaluate loop'.

Add 1-2 more concrete capabilities to the 'what' portion, such as 'runs 5-15 build/test iterations against a weighted rubric (design, originality, craft, functionality) using Playwright' to lift specificity.

DimensionReasoningScore

Specificity

Names the domain and mechanism ('Generator-Evaluator agent harness', 'generator and evaluator iteration until it clears a quality bar') but lists only 1-2 concrete actions, not several specific capabilities. Matches anchor 3: domain plus a couple of concrete actions, not comprehensive; below anchor 4 which expects several listed actions.

3 / 5

Completeness

Both parts present: a clear 'what' ('GAN-inspired Generator-Evaluator agent harness for building high-quality applications autonomously') and an explicit 'Use when a feature should be built autonomously... until it clears a quality bar'. Not 5 because the 'when' clause is abstract ('clears a quality bar') rather than concrete trigger phrases users would actually say.

4 / 5

Trigger Term Quality

Some relevant keywords ('autonomously', 'harness', 'generator and evaluator iteration', 'quality bar') but phrasing is design-paper jargon; natural user terms like 'build an app', 'adversarial loop', 'multi-agent', or 'iterate until it's good' are missing. Fits anchor 3 (relevant keywords, missing common variations) rather than anchor 4's good coverage.

3 / 5

Distinctiveness Conflict Risk

'GAN-inspired', 'Generator-Evaluator', and 'quality bar' carve a distinct niche with low conflict risk against unrelated skills. Not 5 because 'building high-quality applications autonomously' is broad enough to overlap with generic app-building or agentic-development skills.

4 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
affaan-m/ECC
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.