CtrlK
BlogDocsLog inGet started
Tessl Logo

arize-dataset

Creates, manages, and queries Arize datasets and examples. Covers dataset CRUD, appending examples, exporting data, and file-based dataset creation using the ax CLI. Use when the user needs test data, evaluation examples, or mentions create dataset, list datasets, export dataset, append examples, dataset version, golden dataset, or test set.

74

Quality

93%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality CLI reference skill: fully executable commands with real examples, gotchas, and explicit validation feedback loops for batch export/append, with peripheral setup content properly externalized to two well-signaled reference files. Minor improvements possible: reduce repeated '--space' notes and add caution/verification guidance around destructive deletes with --force.

DimensionReasoningScore

Conciseness

Efficient throughout — flag tables and executable commands instead of prose, and the Concepts section covers only Arize-specific product terms. Not 5 because the '--space required when using dataset name instead of ID' note is repeated across ~5 sections and the Related Skills / Save Credentials sections restate peripheral material; not 3 because there is no padding of concepts Claude already knows.

4 / 5

Actionability

Every command is copy-paste ready with concrete flags, real JSON payloads, format gotchas (CSV type loss, JSONL vs JSON array), stdin/heredoc examples, and jq pipelines covering common cases. Not 4 because no key executable detail is missing.

5 / 5

Workflow Clarity

Workflows are clearly sequenced with explicit validation loops — export count verification with re-export '--all' recovery, schema key-check before append, and a verify step after create — satisfying the destructive/batch validation requirement. Not 5 because the destructive 'ax datasets delete --force' path is presented with no caution or target-verification guidance; not 3 because validation checkpoints are present for the batch operations.

4 / 5

Progressive Disclosure

Peripheral troubleshooting/auth content is correctly offloaded to real, one-level-deep bundle files (references/ax-setup.md, references/ax-profiles.md) that are clearly signaled from Prerequisites, the Troubleshooting table, and the Save Credentials section, each scoped with 'Do NOT run proactively'. Core CLI reference appropriately remains inline with per-command section headers. Not 4 because the split and signaling leave no notable organization gaps.

5 / 5

Total

18

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, comprehensive concrete actions, and an explicit 'Use when' clause with natural trigger terms and synonyms. The only weakness is minor trigger overlap with the sibling arize-experiment skill around evaluation/test-data terms.

DimensionReasoningScore

Specificity

Lists multiple concrete actions — 'Creates, manages, and queries Arize datasets and examples', 'dataset CRUD, appending examples, exporting data, and file-based dataset creation using the ax CLI' — comprehensively covering the skill's command surface. Not 4 because there are no minor coverage gaps; not applicable below since every action is concrete rather than generic.

5 / 5

Completeness

Explicitly answers both 'what' ('Creates, manages, and queries Arize datasets and examples... using the ax CLI') and 'when' ('Use when the user needs test data, evaluation examples, or mentions...') with concrete trigger phrases. Not 4 because the 'when' clause is already explicit and specific.

5 / 5

Trigger Term Quality

Natural phrasings users would say include 'create dataset, list datasets, export dataset, append examples, dataset version, golden dataset, or test set' plus 'test data, evaluation examples' — synonyms and variations are covered. Not 4 because no common natural term for this domain is missing.

5 / 5

Distinctiveness Conflict Risk

Clearly niched to Arize datasets via the 'ax CLI' and dataset-specific triggers, but 'test data, evaluation examples' overlaps the closely related arize-experiment skill that consumes these datasets. Not 5 because of that minor overlap with a sibling skill; not 3 because the core triggers are unambiguously dataset-scoped.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
github/awesome-copilot
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.