CtrlK
BlogDocsLog inGet started
Tessl Logo

nemo-curator

GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with RAPIDS. Use for preparing high-quality training datasets, cleaning web data, or deduplicating large corpora.

62

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/ml-training/nemo-curator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

61%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, code-heavy skill that is genuinely actionable for the common curation cases. Its main weaknesses are redundant benchmark/cost sections, missing validation checkpoints in batch curation workflows, and inline multimodal detail that should live in the references directory.

Suggestions

Remove the duplicated performance figures (keep either the GPU vs CPU table or the 'Performance benchmarks' section, not both) and drop or compress the 'Cost comparison' and 'Use cases' marketing sections — they add tokens without instruction.

Add validation checkpoints to the pipeline: e.g., after deduplication report the removed-duplicate fraction, and after filtering check output document counts against expected thresholds before writing to parquet.

Move the image/video/audio curation sections into dedicated reference files (e.g., references/multimodal.md) and link them from the References section, matching how filtering and deduplication are already handled.

DimensionReasoningScore

Conciseness

The body is mostly lean code-plus-headers, but the benchmark numbers appear twice (the 'GPU vs CPU performance' table and the 'Performance benchmarks' section repeat the same 16×/16×/10× figures), and the 'Cost comparison' and 'Use cases / Production deployments' sections are promotional padding rather than instruction. This fits 'mostly efficient but could be tightened' rather than the minor-trim level above.

3 / 5

Actionability

Nearly every section gives concrete, copy-paste-ready Python with real parameters (WordCountFilter(min_words=50...), FuzzyDuplicates(num_hashes=260...)), matching 'mostly executable guidance'. Minor gaps keep it below 5: the multi-GPU/distributed examples use 'FuzzyDuplicates(...)' ellipses, and the Common Crawl pipeline applies filter objects via 'for stage in pipeline: stage(dataset)' when filters elsewhere require dataset.filter(...) — a construct that would not run as written.

4 / 5

Workflow Clarity

The pipeline is clearly sequenced (Stage 1 quality filtering → Stage 2 deduplication → Stage 3 PII redaction → Stage 4 classifier filtering), but there are no validation or verification checkpoints anywhere — no dedup-ratio check, output-count sanity check, or spot-check of filtered data. Per the rubric guideline, batch operations over large corpora without validation cap workflow clarity at 3.

3 / 5

Progressive Disclosure

Structure is good: clear section headers, working code throughout, and a 'References' section linking both real bundle files (references/filtering.md and references/deduplication.md) one level deep with accurate descriptions. It falls short of 5 because substantial detail that belongs in reference files — full image/video/audio curation API examples and cost/benchmark breakdowns — is inlined in SKILL.md while only filtering and deduplication are offloaded.

4 / 5

Total

14

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: comprehensive, third-person, concrete feature list with an explicit 'Use for...' trigger clause. Keyword coverage is good but could add a few natural synonyms, and the web-data-cleaning phrasing slightly overlaps generic data-cleaning skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — 'fuzzy deduplication (16× faster)', 'quality filtering (30+ heuristics)', 'semantic deduplication', 'PII redaction', 'NSFW detection' — covering the toolkit comprehensively. It stays in third person with no vague filler, matching the top anchor.

5 / 5

Completeness

It explicitly answers 'what' ('Features fuzzy deduplication... PII redaction, NSFW detection. Scales across GPUs with RAPIDS.') and 'when' with a concrete trigger clause ('Use for preparing high-quality training datasets, cleaning web data, or deduplicating large corpora'), matching the top anchor exactly.

5 / 5

Trigger Term Quality

Natural phrases like 'cleaning web data', 'deduplicating large corpora', and 'preparing high-quality training datasets' map well to what users would say, and domain terms (deduplication, PII redaction, quality filtering) are present. A few common synonyms such as 'remove duplicates' or 'data cleaning' are missing, so it sits between the 4 and 5 anchors, closer to 4.

4 / 5

Distinctiveness Conflict Risk

The GPU-accelerated LLM-training-data curation niche with multimodal support is clearly distinct, but phrases like 'cleaning web data' create minor overlap risk with generic data-cleaning or datatrove-style skills. This fits the 'mostly distinct; minor overlap risk' anchor rather than the 'minimal conflict risk' anchor above it.

4 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.