CtrlK
BlogDocsLog inGet started
Tessl Logo

training-data-pipeline

Build training datasets for LLM specialization from production data, frontier model distillation, and synthetic bootstrapping. Use when formatting production logs into SFT data, distilling from frontier APIs, or preparing data for fine-tuning. Covers JSONL formatting, data quality validation, deduplication, and train/eval splitting.

62

Quality

74%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/training-data-pipeline/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable — executable code with concrete thresholds, a sequenced checklist, and an error-recovery table — and reasonably lean. Its central weakness is progressive disclosure: a references/ bundle with deeper material exists but is never mentioned, while SKILL.md inlines all path detail itself.

Suggestions

Add explicit links to the existing bundle files, e.g. under each path: 'Deep cleaning and PII redaction: see references/data-quality.md', 'Advanced distillation patterns: see references/frontier-distillation.md', 'Source-specific formatting patterns: see references/production-data-formatting.md'.

Trim the inlined Path A/B/C code listings to minimal examples and move the full implementations into the corresponding reference files to reduce SKILL.md token load.

Fix executable gaps: complete the OpenAI batch submit step, and add 'datasketch' and 'anthropic' to the frontmatter dependencies since the code imports them.

Replace the dated hardcoded model snapshot (claude-sonnet-4-5-20250929) with a current model reference or placeholder to avoid time-sensitive staleness.

DimensionReasoningScore

Conciseness

The body is mostly dense code and tables with little concept-explanation padding, but at ~415 lines it inlines full implementations for all three data paths plus the entire validation/dedup/split toolchain, which could be tightened and partially deferred. Mostly efficient but could be tightened matches the 3 anchor; not 4 because whole sections exceed what the overview role of SKILL.md needs.

3 / 5

Actionability

Nearly all guidance is copy-paste-ready executable Python (api_log_to_training, create_batch_file, validate_jsonl, deduplicate_dataset, split_dataset) with concrete thresholds (0.8 dedup threshold, distinct-2 > 0.5, 90/10 split). Not 5: the OpenAI batch submit step is only a commented-out CLI line, the Anthropic example hardcodes a dated model snapshot, and the code imports 'datasketch' and 'anthropic' which are missing from the frontmatter dependencies.

4 / 5

Workflow Clarity

The Quick Start Checklist gives a clear 8-step sequence with explicit validation checkpoints (validate_jsonl, dedup, diversity report, token count) and the Common Issues table provides error-to-fix feedback loops for this batch-formatting workflow. Not 5: the checklist sits only at the end rather than structuring the document and fixes do not explicitly loop back to re-validation; not 3 because validation steps are explicit, not merely implied.

4 / 5

Progressive Disclosure

Three reference files exist (references/data-quality.md, references/frontier-distillation.md, references/production-data-formatting.md) with deeper complementary material (PII redaction, token length distribution, quality scoring), yet the body never links to any of them, while simultaneously inlining ~400 lines of path-specific detail that belongs at the reference level. Content that clearly belongs in separate files is inlined and the references are buried matches the 2 anchor; not 3 because navigation to the existing bundle is entirely absent, not merely imperfectly signaled.

2 / 5

Total

13

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person concrete capabilities, an explicit 'Use when...' trigger clause, and specific technical terms (SFT, JSONL, deduplication, train/eval split). Keyword coverage misses a few common synonyms, and the fine-tuning-adjacent phrasing creates minor conflict risk with execution skills.

Suggestions

Add common synonyms such as 'supervised fine-tuning (SFT) data', 'instruction tuning', or 'training data preparation' to broaden natural trigger coverage.

Sharpen the boundary with execution skills, e.g. 'use for data preparation only, not for running the fine-tune itself', to reduce overlap with fine-tuning skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'Build training datasets for LLM specialization from production data, frontier model distillation, and synthetic bootstrapping' plus 'JSONL formatting, data quality validation, deduplication, and train/eval splitting' — covering all major sub-tasks. Not 4: coverage is comprehensive with no meaningful gaps, matching the top anchor.

5 / 5

Completeness

It explicitly answers both what ('Build training datasets... Covers JSONL formatting, data quality validation, deduplication, and train/eval splitting') and when ('Use when formatting production logs into SFT data, distilling from frontier APIs, or preparing data for fine-tuning') with concrete trigger phrases. Not 4: the 'when' clause is specific and actionable, not merely present.

5 / 5

Trigger Term Quality

Natural user phrases are present ('formatting production logs into SFT data', 'distilling from frontier APIs', 'preparing data for fine-tuning', 'deduplication'), but synonyms like 'instruction tuning' or 'training data preparation' are absent. Good coverage with a few natural terms missing fits the 4 anchor; 5 would require synonym-level breadth.

4 / 5

Distinctiveness Conflict Risk

The training-data-preparation niche is clear, but 'preparing data for fine-tuning' and 'distilling from frontier APIs' create minor overlap risk with fine-tuning-execution skills (the body itself defers to 'tinker'/'unsloth' skills). Mostly distinct with minor overlap against closely related skills fits the 4 anchor; 5 would require unambiguous triggers against all neighbors.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.