CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-reproduce-align

Use after a Codex or Claude Code feature has been implemented in Qwen Code to run the selected reference agent and Qwen Code under the same scenario, capture HTTP and terminal traces, compare request bodies, tool/function schemas, outputs, and iterate until the reproduced behavior is close enough.

75

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality skill body: concise, immediately actionable commands, a well-sequenced workflow with an explicit iterate-until-done loop, and clean separation of overview versus reference detail. The only gap is that a couple of execution variants (Claude Code command discovery, manual tmux capture) are described rather than shown.

DimensionReasoningScore

Conciseness

The body is lean and assumes competence: no concept explanations, no padding — every section (selection rule, 8-step workflow, three command blocks, comparison rules, done criteria) carries operational content. It matches the 5 anchor; nothing reads as over-explanation that would justify 4.

5 / 5

Actionability

Three copy-paste-ready shell blocks (normalize, compare, paired capture with env vars) plus an ordered diff-inspection checklist make the guidance mostly executable. It falls short of 5 only because some paths are left to discovery ('replace the first command with the discovered Claude Code command', run captures 'manually with tmux' without a concrete invocation).

4 / 5

Workflow Clarity

The 8-step workflow is clearly sequenced and includes an explicit feedback loop ('Patch Qwen Code, rerun the smallest failing scenario, and repeat') plus a Done Criteria checklist that serves as validation checkpoints, with further iteration detail in the reference file. This matches the 5 anchor; the operation is not destructive/batch so no cap applies.

5 / 5

Progressive Disclosure

The body is a well-organized overview with one clearly signaled, one-level-deep reference ('Read references/alignment-workflow.md before the first comparison pass' — a real file) and three real scripts referenced by name, with detailed diff-triage priorities appropriately split into the reference. Structure matches the 5 anchor with easy navigation.

5 / 5

Total

19

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete capabilities, gives an explicit use-after trigger tied to a clear workflow position, and occupies a distinctive niche. The only weakness is modest keyword variety — natural synonyms like 'parity' or 'behavior matching' that a user might say are absent.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'run the selected reference agent and Qwen Code under the same scenario, capture HTTP and terminal traces, compare request bodies, tool/function schemas, outputs, and iterate' — covering the skill's full scope with no generic filler. It matches the 5 anchor (comprehensive list of specific concrete actions) rather than 4, since there are no meaningful coverage gaps.

5 / 5

Completeness

Both questions are answered explicitly: 'what' via the concrete action list (capture, compare, iterate) and 'when' via the explicit trigger 'Use after a Codex or Claude Code feature has been implemented in Qwen Code'. This matches the 5 anchor with a concrete trigger phrase; it is not 4 because the 'when' clause is already explicit and specific.

5 / 5

Trigger Term Quality

Strong domain keywords users would naturally say: 'Codex', 'Claude Code', 'Qwen Code', 'reference agent', 'traces', 'compare', 'reproduce'. A few natural synonyms are missing (e.g., 'parity', 'matching behavior', 'diff traces'), so it sits between the 4 and 5 anchors but clearly above 3, where common variations would be absent.

4 / 5

Distinctiveness Conflict Risk

The Qwen Code reference-agent parity niche is highly specific and the trigger conditions (feature implemented, parity evidence needed) would not fire for unrelated skills. Minimal conflict risk, matching the 5 anchor's 'clear niche with distinct triggers'.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
QwenLM/qwen-code
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.