CtrlK
BlogDocsLog inGet started
Tessl Logo

trace-to-training-data

Convert evaluation traces and production logs into SFT examples and preference pairs. Use when graded traces or failure examples exist and need to become training data, when applying rejection sampling to model outputs, or when building DPO pairs from passing and failing runs.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-built overview skill: principles and concrete thresholds live in the body, executable recipes live in a single verified one-level-deep reference, and the hygiene section supplies real fail-closed validation for a batch operation. The main improvement levers are tightening repeated rationale in the body and specifying the couple of deferred details (rejection-sampling fraction, secret-scan method) or explicitly routing them to the reference.

Suggestions

State the default rejection-sampling keep fraction (e.g. the 0.25 used in references/conversion-recipes.md) in the SFT section, or explicitly defer it to the recipe the way the μ−2σ formula is deferred to preference-optimization.

Name a concrete secret/PII scanning approach (regex patterns, a specific scanner) in the Hygiene section so 'run a secret/PII scan' is executable rather than aspirational.

Trim the duplicated goldens-leakage rationale — it is explained at equal length in both the Hygiene section and the reference's section 5.

DimensionReasoningScore

Conciseness

The body is efficient — it assumes the flywheel context and never re-explains SFT/DPO basics, e.g. 'Curation is the work that remains — which traces clear a quality bar...'. Not 5 because some rationale is repeated (the goldens-leakage warning appears in full in both the body and the reference) and benchmark citations like 'SRFT reports 32.2% vs. 30.9% on SWE-bench' pad beyond what the instruction needs; not 3 because the padding is minor, not a section's worth of over-explanation.

4 / 5

Actionability

Concrete numeric guidance ('select the rejected member at μ−2σ', 'the top 5k of a 16.5k candidate pool'), an explicit input JSON schema, and exact field mappings, with fully executable worked conversions verified to live in references/conversion-recipes.md (select_pair, rejection_sample, filter_holdout code). Not 5 because some body guidance defers without specifying — the top 'fraction' for rejection sampling is unnamed in the body and the secret/PII scan names no tool or pattern; not 3 because the bundle as a whole gives executable, copy-paste-ready recipes for the common cases.

4 / 5

Workflow Clarity

Validation checkpoints for this batch operation are explicit and fail-closed: 'conversion fails closed (the row is dropped...)', the goldens holdout, dedup against the existing training set, and the provenance check, plus 'Only when valid' style gating in the reference's holdout recipe ('Run filter_holdout before the messages-only stripping'). Not 5 because the sequence itself is implicit — the workflow must be assembled from the SFT, Pairs, and Hygiene sections rather than a numbered flow; not 3 because checkpoints are present and explicit, which the score-3 anchor lacks.

4 / 5

Progressive Disclosure

The body is a well-sectioned overview (Principle, SFT, Pairs, Hygiene, Related Skills) and defers all worked examples to a single reference that exists, is one level deep, and whose contents are enumerated at the point of reference ('a graded trace to SFT row, trace pair to DPO pair, correction to SFT row, the rejection-sampling loop, and the goldens-holdout check... live in references/conversion-recipes.md'). Verified the reference file matches that description exactly with no nested references, matching the score-5 anchor.

5 / 5

Total

17

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a strong example: it states concrete capabilities in third person and gives an explicit multi-clause 'Use when...' with natural trigger phrases covering rejection sampling and DPO pair construction. Its only weaknesses are minor — a few missing synonyms and no mention of the hygiene steps the body treats as core work.

DimensionReasoningScore

Specificity

Lists several specific actions — 'Convert evaluation traces and production logs into SFT examples and preference pairs', 'applying rejection sampling', 'building DPO pairs from passing and failing runs' — which goes beyond the 1-2 actions of a score-3 anchor. Not 5 because the skill's hygiene duties (secret scanning, goldens holdout, provenance) are absent from the description, leaving a minor coverage gap.

4 / 5

Completeness

Explicitly answers both: 'what' — converting traces/logs into SFT examples and preference pairs — and 'when' via a three-clause 'Use when...' ('when graded traces or failure examples exist and need to become training data, when applying rejection sampling..., or when building DPO pairs from passing and failing runs') with concrete trigger phrases. Matches the score-5 anchor exactly.

5 / 5

Trigger Term Quality

Natural practitioner phrases like 'training data', 'rejection sampling', 'DPO pairs', and 'passing and failing runs' give good keyword coverage a user would plausibly say. Not 5 because common synonyms such as 'fine-tuning dataset' or 'preference data' are missing; not 3 because multiple natural multi-word trigger phrases are present, not just a generic keyword or two.

4 / 5

Distinctiveness Conflict Risk

Clear niche — converting already-graded eval traces into training data — with distinct triggers ('graded traces', 'rejection sampling to model outputs', 'DPO pairs from passing and failing runs') that would not naturally fire for sibling skills like dataset curation or eval harnessing. Minimal conflict risk, matching the score-5 anchor; the score-4 anchor's 'minor overlap risk' does not fit since the trigger terms are unique to this conversion step.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
wshobson/agents
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.