CtrlK
BlogDocsLog inGet started
Tessl Logo

dataset-discovery

Multi-source ML dataset discovery. Search HuggingFace Hub, OpenML, GitHub, and paper cross-references for datasets relevant to a research task. Use when asked to "find datasets for", "search ML datasets", "what datasets exist for", or "discover training data for".

68

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable skill body: every phase is backed by executable commands verified against the real script, and the single script dependency is correctly disclosed. Weaknesses are minor — a redundant Overview paragraph and the absence of empty-results/error-handling guidance in the search workflow.

DimensionReasoningScore

Conciseness

The body is lean and command-first, but the Overview paragraph largely restates the frontmatter description (with a naming inconsistency: 'Semantic Scholar' vs 'paper cross-references'), and dependency notes like 'stdlib-adjacent, universally available' could be trimmed — matching anchor 4's 'minor instances of over-explanation'.

4 / 5

Actionability

Every phase gives copy-paste-ready commands whose subcommands, flags, and defaults match the actual script CLI, plus a concrete example dataset id ('huggingface:stanfordnlp/imdb'), specified output paths, and a defined presentation table — fully executable and covering the common cases, matching anchor 5.

5 / 5

Workflow Clarity

The five phases are clearly sequenced with per-phase commands and a user-confirmation checkpoint before full downloads, but there are no checkpoints for empty search results or per-source API failures — anchor 4's 'clear sequence with most checkpoints; minor validation gaps' fits. The destructive/batch cap does not apply since all operations are read-only searches.

4 / 5

Progressive Disclosure

All logic lives in the verified one-level-deep script reference (scripts/search_ml_datasets.py), and the ~80-line body is a well-organized overview with no content that belongs in separate reference files — appropriately split and easy to navigate, matching anchor 5.

5 / 5

Total

18

/

20

Passed

Description

78%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with an explicit trigger clause, natural quoted trigger phrases, and a concrete, well-scoped niche. Its main weakness is that it describes a single search action rather than enumerating several concrete capabilities (e.g., detail retrieval, sample pulls, ranked dedup).

Suggestions

Add one or two more concrete action verbs to the description (e.g., 'fetch detailed metadata and pull sample rows, returning a ranked, deduplicated list') to lift specificity from one search action to several capabilities.

Include a couple of natural trigger synonyms such as 'corpus' or 'benchmark datasets' to broaden keyword coverage toward the top anchor.

DimensionReasoningScore

Specificity

The description names the domain and sources concretely ('Search HuggingFace Hub, OpenML, GitHub, and paper cross-references for datasets') but relies on essentially one action verb (search/discovery), matching anchor 3's '1-2 concrete actions, not comprehensive' rather than anchor 4's 'several specific actions'.

3 / 5

Completeness

It clearly answers 'what' (multi-source search across named sources for datasets relevant to a research task) and explicitly answers 'when' with concrete trigger phrases via the 'Use when asked to...' clause, matching the anchor-5 example pattern exactly.

5 / 5

Trigger Term Quality

Four natural quoted trigger phrases ('find datasets for', 'search ML datasets', 'what datasets exist for', 'discover training data for') give good keyword coverage, but common variations like 'corpus', 'benchmark datasets', or modality-specific phrasing are missing, so it does not reach anchor 5's comprehensive synonym coverage.

4 / 5

Distinctiveness Conflict Risk

The ML-dataset-discovery niche with named sources and distinctive triggers is well-differentiated, but there is minor overlap risk with generic web-search or HuggingFace-specific search skills, so anchor 4 ('mostly distinct; minor overlap risk') fits better than anchor 5.

4 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
OpenLAIR/dr-claw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.