CtrlK
BlogDocsLog inGet started
Tessl Logo

dataset-curation

Prepare, format, and validate datasets for supervised fine-tuning and preference training. Use when converting raw data into training format, applying chat templates, configuring sequence packing, generating synthetic training data, or writing a dataset card before a run.

71

Quality

86%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

85%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable body with strong workflow gating and clean progressive disclosure into two real reference files. Tightening the few explanatory passages and inlining a touch more executable code would raise conciseness and actionability.

DimensionReasoningScore

Conciseness

Mostly lean and operational with terse prose and concrete thresholds, but a few explanatory passages (e.g., the padding rationale and synthetic-data collapse justification) could be tightened without losing clarity.

4 / 5

Actionability

Provides executable code snippets (loss-mask decode check, packed-sequence inspection), a format-selection table, and concrete numeric thresholds, but defers fuller code sketches to the reference files, leaving minor gaps in inline coverage.

4 / 5

Workflow Clarity

Clear sequence from format selection through templating, packing, synthetic-data mixing, and the dataset card, with explicit validation checkpoints ('MANDATORY: decode and manually inspect 5–10 packed sequences') and a six-item Phase 2 Exit Checklist acting as a gate.

5 / 5

Progressive Disclosure

SKILL.md is a well-organized overview with clearly signaled, one-level-deep references to two real files (formats-and-templates.md, synthetic-data.md) cited inline and re-listed at the end, with bulk detail appropriately split out.

5 / 5

Total

18

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly communicates a distinct niche and concrete capabilities with an explicit 'Use when' trigger clause. Adding common abbreviations (SFT, DPO, JSONL) as synonyms would push trigger-term quality to full marks.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Prepare, format, and validate datasets', 'converting raw data into training format, 'applying chat templates', 'configuring sequence packing', 'generating synthetic training data', 'writing a dataset card' — giving comprehensive coverage of the skill's capabilities.

5 / 5

Completeness

Explicitly states both what ('Prepare, format, and validate datasets for supervised fine-tuning and preference training') and when ('Use when converting raw data into training format, ... or writing a dataset card before a run') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Strong natural domain terms ('supervised fine-tuning', 'preference training', 'chat templates', 'sequence packing', 'synthetic training data', 'dataset card') a practitioner would say, but lacks common synonyms/abbreviations like SFT, DPO, or JSONL that users also employ.

4 / 5

Distinctiveness Conflict Risk

The dataset-curation niche is clearly distinct with specific triggers, but it sits in a dense finetuning-skill ecosystem (method-selection, preference-optimization) where minor overlap risk exists for closely related skills.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
wshobson/agents
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.