CtrlK
BlogDocsLog inGet started
Tessl Logo

nemo-curator

Curate LLM training data: dedupe, filter, PII redaction.

48

Quality

53%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/nemo-curator/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

53%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured and code-forward with a commendably honest 1.x/0.x migration warning, but it suffers from duplicated benchmark content, inline code that repeats what the reference files already cover, and no validation checkpoints for destructive batch operations. Much of the code is self-admittedly non-runnable in the current version, which limits its actionability despite concrete installation steps.

Suggestions

Add validation checkpoints to the curation pipeline (e.g. 'Tune thresholds on a 10k-doc sample before full runs' and 'Verify output row counts and spot-check redacted samples before overwriting'), which would raise workflow clarity above the batch-operation cap of 3.

Remove the duplicated performance numbers by keeping either the GPU-vs-CPU table or the 'Performance benchmarks' section, not both, and drop the cost-comparison detail to a single summary line.

De-duplicate the Stage 1/Stage 2 filter and dedup code by pointing to references/filtering.md and references/deduplication.md, reserving the body for the pipeline shape and one minimal end-to-end 1.x example.

Since 0.x snippets are flagged as conceptual, replace the longest 0.x sections with the actual 1.x stage composition so the code is copy-paste ready.

DimensionReasoningScore

Conciseness

Prose is lean and code-dense, but the GPU-vs-CPU numbers appear twice (the table under 'GPU acceleration' and again under 'Performance benchmarks'), and the Stage 1/Stage 2 code substantially duplicates content already in references/filtering.md and references/deduplication.md. Mostly efficient but could be tightened, matching the 3 anchor.

3 / 5

Actionability

Installation commands are copy-paste ready, but the skill itself flags the bulk of its code as 'conceptual (0.x-style)' with 'treat the examples in this skill below as conceptual', and snippets like 'FuzzyDuplicates(...)' are placeholders. The deprecation is explicitly acknowledged and linked to the 1.x quickstart, which keeps it from falling to 2, but the guidance is not copy-paste executable, matching the 3 anchor.

3 / 5

Workflow Clarity

The curation pipeline is clearly sequenced (install -> quality filtering -> dedup -> PII redaction -> classifier filtering) and the 1.x pipeline shape is shown, but there are no validation or verification checkpoints for destructive batch operations on multi-TB datasets, which caps workflow clarity at 3 per the judging guidelines.

3 / 5

Progressive Disclosure

Both referenced files (references/filtering.md, references/deduplication.md) exist, are one level deep, and are clearly signaled in a 'References' section with descriptions; section structure is good. Not 5 because the inline Stage 1/2 code duplicates the reference-file content that should have been kept separate; not 3 because references are clearly signaled and most content (multi-modal, GPU, cost) is appropriately placed.

4 / 5

Total

13

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concise and names a distinct niche with three actions, but it is terse at the cost of completeness and trigger coverage: it lacks any 'Use when' trigger guidance and misses common synonyms like 'deduplication' or 'data cleaning'. Adding a when-to-use clause and a few natural trigger phrases would lift it substantially.

Suggestions

Add a 'Use when...' clause, e.g. 'Use when preparing LLM training data from web scrapes or when the user mentions deduplication, data cleaning, PII redaction, or curating datasets.'

Include natural synonyms such as 'deduplication' and 'data cleaning/dataset preparation' so the description matches how users actually phrase these needs.

Name one or two distinctive capabilities (e.g. GPU-accelerated fuzzy dedup, multi-modal curation) to sharpen specificity beyond the generic word 'filter'.

DimensionReasoningScore

Specificity

Names the domain ('LLM training data') and lists three actions ('dedupe, filter, PII redaction'), but 'filter' is generic and coverage is not comprehensive (no mention of multi-modal curation, classifiers, or data formats). It sits above the 3 anchor in action count but not enough to reach 4's 'several specific actions' with only minor gaps.

3 / 5

Completeness

A clear 'what' is present ('Curate LLM training data: dedupe, filter, PII redaction') but there is no 'Use when...' clause or equivalent trigger guidance, which caps completeness at 3 per the judging guidelines. It is not 2 because the 'what' is concrete and specific.

3 / 5

Trigger Term Quality

'dedupe', 'PII redaction', and 'LLM training data' are natural user phrases, but common variations and synonyms ('deduplication', 'data cleaning', 'dataset preparation', 'training corpus') are missing. Not 4 because keyword coverage is thin for a nine-word description.

3 / 5

Distinctiveness Conflict Risk

The 'LLM training data' curation niche with dedupe/PII-redaction triggers is mostly distinct, with only minor overlap risk against general data-processing skills. Not 5 because 'dedupe' and 'filter' alone could also match generic data-cleaning tooling.

4 / 5

Total

13

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.