CtrlK
BlogDocsLog inGet started
Tessl Logo

agent-reproduce-align

Use after a Codex or Claude Code feature has been implemented in Qwen Code to run the selected reference agent and Qwen Code under the same scenario, capture HTTP and terminal traces, compare request bodies, tool/function schemas, outputs, and iterate until the reproduced behavior is close enough.

70

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable skill body with executable commands, a clear sequenced workflow, and good progressive disclosure via a single reference file and bundled scripts. Minor tightening and a more concrete Claude Code command path would lift actionability and conciseness.

Suggestions

Provide an explicit, copy-pasteable command template for the Claude Code reference agent rather than 'replace the first command with the discovered Claude Code command'.

Add an explicit inline validation gate in the workflow (e.g., 'only proceed to patching when compare_traces.py reports no must-match failures') to strengthen the feedback loop.

Tighten a few prose sentences in Comparison Rules and the paired-runner paragraph to reduce token cost without losing clarity.

DimensionReasoningScore

Conciseness

Lean and efficient with no concept-explanation fluff; sections like Comparison Rules and Done Criteria earn their place, though a few sentences could be tightened slightly, keeping it just below the 5 anchor.

4 / 5

Actionability

Provides copy-paste ready shell commands for normalize, compare, and paired runs with real paths and arguments; the Claude Code variant ('replace the first command with the discovered Claude Code command') is a minor gap versus fully executable.

4 / 5

Workflow Clarity

An 8-step sequenced workflow with a rerun-and-repeat feedback loop (step 7) and a Done Criteria checklist; validation is present as a retry loop and checklist rather than an explicit inline gate, so it sits below the 5 anchor.

4 / 5

Progressive Disclosure

Well-organized overview sections with one clearly signaled one-level-deep reference ('Read references/alignment-workflow.md before the first comparison pass') and bundle scripts referenced appropriately, matching the 5 anchor.

5 / 5

Total

17

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description with concrete actions, an explicit trigger clause, and a distinct niche. Slightly room for more natural synonyms like 'parity' or 'reproduce' in the trigger terms.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'run the selected reference agent and Qwen Code under the same scenario, capture HTTP and terminal traces, compare request bodies, tool/function schemas, outputs, and iterate' — giving comprehensive coverage comparable to the 5 anchor.

5 / 5

Completeness

Explicitly answers both what (run, capture, compare, iterate) and when ('Use after a Codex or Claude Code feature has been implemented in Qwen Code'), with a concrete trigger phrase, matching the 5 anchor.

5 / 5

Trigger Term Quality

Good keyword coverage including both reference agents ('Codex or Claude Code'), 'Qwen Code', 'traces', 'request bodies', 'tool/function schemas', 'outputs', but a few natural terms like 'parity' or 'reproduce' are absent, keeping it just below comprehensive.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (parity reproduction between Qwen Code and Codex/Claude Code) with distinct triggers and minimal overlap risk with other skills.

5 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
QwenLM/qwen-code
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.