CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-release-gate

Run the agent release gate — a portable, wire-level QA harness for the agent runtime. Drives the same product endpoint the playground drives and asserts on the SSE frame stream and real side effects, never on model prose, so it works against any deployment (cloud or self-hosted) from three env vars. Use before an agent-workflows release, or after changing the runner, the SDK agent adapters, the runner Docker images, or the agent service. Triggers: "run the release gate", "QA the agent runtime", "does the agent still work end to end", "pre-release agent QA".

68

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is exceptionally actionable and validation-dense — copy-paste commands, exact flags and defaults, and an explicit checkpoint on every risky step — but it badly overruns the token budget: roughly 700 lines where the dated incident narratives, session IDs, per-cell evidence history, and version-sensitive notes belong in the reference files the skill itself points to (coverage.md, LESSONS.md), which are not even shipped in this bundle.

Suggestions

Conciseness: move the dated incident narratives and verification records (the 2026-08-06 PASS histories, session UUIDs, the v0.117.0 preview-stage findings) into resources/LESSONS.md or an evidence file, keeping one-line takeaways inline — the body repeatedly cites these as 'read on demand' material while inlining them anyway.

Progressive disclosure: collapse each matrix cell entry in the Resources section to one line (name, tier, what it proves, what it needs) and leave the scenario detail and gotchas to the files' own docstrings and resources/coverage.md, so SKILL.md is an overview rather than a per-cell runbook.

Conciseness: gather version- and date-sensitive guidance ('if you are reading an old green from before 2026-08-06', the v0.117.0 runner /run contract changes, the v0.115.3 gate gap) into a clearly marked old-patterns/changelog reference so time-sensitive material stops padding the mainline instructions.

DimensionReasoningScore

Conciseness

Noticeably verbose with several padded sections: dated incident narratives ('Cost a staging gate run on 2026-08-28', 'left the runner down for about seven minutes mid-gate on 2026-09-10'), raw session UUIDs ('58ce3a58-8d04-40ac-99e4-c44eaa5d7b06'), and open working notes ('Worth a call: whether the guidance text needs to be more directive') fill a ~700-line body. Time-sensitive dates and versions (v0.117.0, 2026-08-06, v0.115.3) sit inline rather than in an old-patterns/lessons file. Not 1: it does not explain concepts Claude already knows — the padding is domain-specific narrative, not tutorial filler, and the core instructions are dense with signal.

2 / 5

Actionability

Fully executable throughout: the three env vars with export discipline, copy-paste `uv run resources/qa_product.py --all --custom-slug <vault-slug> ...` invocations with every flag and default enumerated, and a complete SIGKILL/start/health-wait hook script (heredoc with bounded loop and non-zero exit on timeout). Commands cover the common cases from full run down to a single journey. Not 4: there is no missing key detail — even failure modes name the exact error strings.

5 / 5

Workflow Clarity

Clear sequence with explicit validation at every risky step: 'confirm the stage in the results before trusting them', --require-store forcing continuity journeys to FAIL rather than silently SKIP, 'Any FAIL blocks the release until triaged', 'read resources/LESSONS.md' before trusting a pass, SKIPs named as untested claims in the summary, and --session-control-results stopping the gate on a missing artifact. Batch and destructive operations (16-run bursts, SIGKILL of the runner replica) each carry feedback loops (bounded health wait with status output, provider-capacity SKIP handling, resume from prior results.json). Not 4: checkpoints are present at every step, not just most.

5 / 5

Progressive Disclosure

A well-signaled 'Resources (read on demand)' section with one-line meanings per file exists, but the body inlines what it delegates: per-cell runbooks with dated verification history (matrix_w7, matrix_l1..l5, matrix_b1 each get a paragraph of PASS/session evidence), the full preview-stage section, and the runner /run contract — material its own pointers say lives in resources/coverage.md and LESSONS.md. The bundle ships none of the ~25 referenced resources/ files, so no referenced path resolves in this skill bundle. Not 2: structure is real (clear headers, an indexed resource list), and references are clearly signaled, not buried.

3 / 5

Total

15

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: it states what the harness does at the mechanism level (wire-level assertions on SSE frames and real side effects, never model prose), its deployment-agnostic scope, exactly when to invoke it, and a natural trigger-phrase list. All four dimensions sit at the top anchor; no suggestions.

DimensionReasoningScore

Specificity

Multiple concrete actions with comprehensive coverage: 'Run the agent release gate — a portable, wire-level QA harness', 'Drives the same product endpoint the playground drives and asserts on the SSE frame stream and real side effects, never on model prose', 'works against any deployment (cloud or self-hosted) from three env vars'. Every verb names a specific, verifiable action. Not 4: anchor 4 allows 'minor gaps in coverage', but what the skill does, how it asserts, and its deployment scope are all explicitly stated with no gap.

5 / 5

Completeness

Explicitly answers both questions. What: a wire-level QA harness that drives the product endpoint and asserts on SSE frames and real side effects, never model prose. When: 'Use before an agent-workflows release, or after changing the runner, the SDK agent adapters, the runner Docker images, or the agent service', plus concrete trigger phrases. Not 4: the 'when' is not merely present but specific and multi-conditioned, matching the anchor-5 example's structure.

5 / 5

Trigger Term Quality

Four natural trigger phrases a user would actually say ('run the release gate', 'QA the agent runtime', 'does the agent still work end to end', 'pre-release agent QA') plus context triggers ('after changing the runner, the SDK agent adapters, the runner Docker images, or the agent service'). Not 4: no common phrasing of this need is missing — the explicit 'Triggers:' list covers synonyms across command, question, and noun forms.

5 / 5

Distinctiveness Conflict Risk

Clear niche — agent-runtime release QA over the wire — with distinct, unambiguous triggers ('agent release gate', 'pre-release agent QA') that no other plausible skill answers. Minimal overlap risk; not 4 because no closely related skill would capture these triggers first.

5 / 5

Total

20

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (713 lines); consider splitting into references/ and linking

Warning

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

Total

14

/

16

Passed

Repository
Agenta-AI/agenta
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.