CtrlK
BlogDocsLog inGet started
Tessl Logo

nemo-curator

GPU-accelerated data curation for LLM training. Supports text/image/video/audio. Features fuzzy deduplication (16× faster), quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection. Scales across GPUs with RAPIDS. Use for preparing high-quality training datasets, cleaning web data, or deduplicating large corpora.

66

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

68%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Well-structured and actionable content with real progressive-disclosure references, but it lacks validation/verification checkpoints for batch curation operations and carries some redundant benchmark restatements. Adding error-recovery feedback loops would materially improve it.

Suggestions

Add explicit validation checkpoints to the curation pipeline (e.g., verify record counts between stages, validate parquet schema after dedup, and a re-run-on-failure step) so batch operations have a feedback loop.

Consolidate the repeated 16× performance figures into one benchmark section and reference it, removing the duplicated numbers from Quick start and the GPU table to tighten conciseness.

Move the image/video/audio curation blocks into separate reference files (e.g., references/multimodal.md) and link from the body, mirroring the existing filtering/deduplication pattern.

DimensionReasoningScore

Conciseness

Mostly lean with direct headers and code blocks, but performance figures (16×) are restated across Quick start, GPU table, benchmarks, and cost sections, and the cost-comparison block adds length that could be trimmed.

4 / 5

Actionability

Provides concrete, executable installation commands and parameterized Python snippets, but a few scaling examples use FuzzyDuplicates(...) placeholders, leaving minor gaps.

4 / 5

Workflow Clarity

The pipeline is clearly staged (filter -> dedup -> PII -> classifier), but batch operations over large corpora run with no validation or verification checkpoints, which caps this dimension at 3 per the rubric.

3 / 5

Progressive Disclosure

Two real, clearly signaled one-level-deep references (filtering.md, deduplication.md) are organized in a References section and exist on disk; minor gap is that image/video/audio content stays inline rather than being split out.

4 / 5

Total

15

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong, specific description that clearly states capabilities and explicit use-trigger phrases, with a distinct niche. Keyword coverage is good but could add common synonyms and file extensions to be fully comprehensive.

DimensionReasoningScore

Specificity

Names the domain and lists multiple concrete actions — fuzzy deduplication, quality filtering (30+ heuristics), semantic deduplication, PII redaction, NSFW detection — with comprehensive coverage matching the score-5 anchor.

5 / 5

Completeness

Explicitly answers both what (curation, dedup, filtering, PII, NSFW) and when ('Use for preparing high-quality training datasets, cleaning web data, or deduplicating large corpora') with concrete trigger phrases.

5 / 5

Trigger Term Quality

Includes natural trigger phrases ('preparing high-quality training datasets', 'cleaning web data', 'deduplicating large corpora') with good coverage, but misses common synonyms and file-format terms that would lift it to 5.

4 / 5

Distinctiveness Conflict Risk

Occupies a clear niche (GPU-accelerated LLM data curation via RAPIDS, multimodal) with distinct triggers like 'NeMo Curator' and 'GPU acceleration', giving minimal conflict risk.

5 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.