CtrlK
BlogDocsLog inGet started
Tessl Logo

zarr-python

Chunked N-D arrays for cloud storage. Compressed arrays, parallel I/O, S3/GCS integration, NumPy/Dask/Xarray compatible, for large-scale scientific computing pipelines.

53

Quality

61%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/data-engineering/zarr-python/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A thorough, code-heavy reference that is highly actionable across the full Zarr feature surface, held back by noticeable verbosity (repeated chunk-size/metadata guidance across four-plus sections) and a progressive-disclosure defect: the bundled references/api_reference.md is never linked from SKILL.md, so it is effectively undiscoverable. Trimming duplication and splitting patterns/API detail into the existing reference file would materially improve it.

Suggestions

Link references/api_reference.md from the body (e.g., '## API reference — see [api_reference.md](references/api_reference.md)') so the bundled file is actually discoverable, and move the compression codec details and 'Common Patterns' walkthroughs there.

Deduplicate the chunk-size guidance (currently repeated in Chunking Strategies, the Performance checklist, the Cloud-Native pattern, and Common Issues) into a single canonical section.

Fix minor executability gaps: define `data`/`data2`, import pandas in the Xarray example, and separate or annotate v2 vs v3 API usage (e.g., S3Map + create_array).

DimensionReasoningScore

Conciseness

The body is noticeably verbose at ~770 lines: chunk-size guidance is repeated in four places (Chunking Strategies, the Performance Optimization checklist, the Cloud-Native pattern, and Common Issues), consolidated-metadata advice appears in both 'Cloud Storage Best Practices' and its own section, and there are several padded 'Benefits' bullet lists and explanatory asides (e.g., 'Groups organize multiple arrays hierarchically, similar to directories or HDF5 groups') that restate knowledge Claude already has. It is not a score-1 tutorial-for-beginners dump — most content is code, not prose — but the duplication and padding clearly sit below the 'mostly efficient' score-3 anchor.

2 / 5

Actionability

Nearly all guidance is concrete, executable code covering creation, indexing, chunking, compression, storage backends, cloud I/O, groups, Dask/Xarray integration, and diagnosis of common issues (e.g., 'print(z.chunks)' followed by specific solutions). Minor gaps keep it from a 5: some examples use undefined variables (`data` in the Cloud-Native pattern, `data2` and an unimported `pd` in the Xarray example), and the API mixes v2-style (`s3fs.S3Map`, `zarr.append`) with v3-style (`zarr.create_array`, `shards=`) calls that may not run together as written.

4 / 5

Workflow Clarity

For a library/reference skill, the body provides a clear decision sequence — a numbered 'Checklist for Optimal Performance' (chunk size → chunk shape → compression → backend → sharding → parallel I/O) plus 'Common Issues' sections structured as diagnosis code → concrete solutions, which function as feedback checkpoints. It falls short of 5 because there is no explicit end-to-end workflow with validation/error-recovery loops, and a few sequences (e.g., the Cloud-Native pattern's write-then-consolidate flow) omit checkpoints like verifying the write or handling partial failures.

4 / 5

Progressive Disclosure

Section headers are well-organized and one-level-deep, but the bundle contains references/api_reference.md (a 12KB function reference) that is never mentioned or linked anywhere in SKILL.md — an agent following the skill would never discover it. Meanwhile large blocks of inline content (compression codec tables, the four 'Pattern' walkthroughs, profiling details) belong in reference files per the score-3/4 anchors. This matches 'references present but not clearly signaled; content that should be separate is inline'; it is above score 2 because the document is well structured, not a minimal-structure wall of text.

3 / 5

Total

13

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A specific, third-person description with good keyword coverage, but it omits any explicit 'Use when…' trigger guidance and — surprisingly for a library skill — never mentions the library's own name 'Zarr' or the '.zarr' extension, which weakens both completeness and trigger-term quality. Solid but below the exemplary anchor examples.

Suggestions

Add an explicit trigger clause, e.g.: 'Use when working with Zarr stores, chunked arrays, .zarr files, or large scientific datasets that don't fit in memory.'

Include the library name and file extension in the description text itself ('Zarr', '.zarr'), since users will say 'zarr' when they need this skill.

State capabilities as concrete actions ('Create, read, and append to chunked arrays…') rather than a noun-phrase feature list.

DimensionReasoningScore

Specificity

The description lists several concrete capabilities — 'Compressed arrays, parallel I/O, S3/GCS integration, NumPy/Dask/Xarray compatible' — in third person, with only minor gaps (no mention of chunking strategies, groups/hierarchies, or Dask/Xarray write workflows). It lists several specific capabilities rather than only 1-2, placing it above the score-3 anchor, but the noun-phrase feature list ('parallel I/O') is less action-verb concrete than the score-5 anchor's 'Extract text…, fill forms…, merge documents'.

4 / 5

Completeness

The 'what' is clear ('Chunked N-D arrays for cloud storage. Compressed arrays, parallel I/O, S3/GCS integration…') but there is no 'Use when…' clause or equivalent explicit trigger guidance — 'for large-scale scientific computing pipelines' only weakly implies when. Per the judging guideline, a missing explicit 'when' clause caps completeness at 3 even though the 'what' is solid, so it cannot score 4.

3 / 5

Trigger Term Quality

Good natural keyword coverage: 'N-D arrays', 'cloud storage', 'S3/GCS', 'NumPy/Dask/Xarray', 'scientific computing' are phrases users would naturally say. It misses the most important trigger terms — the description never says 'Zarr' (only the name field does), nor the '.zarr' extension or synonyms like 'chunked storage', so it does not reach the score-5 anchor requiring synonyms and file extensions.

4 / 5

Distinctiveness Conflict Risk

The chunked-array/cloud-storage niche is mostly distinct with concrete signals (S3/GCS, Dask, Xarray, scientific computing), so it is above the 'somewhat specific but could overlap' anchor. Minor overlap risk remains: mentioning 'NumPy/Dask/Xarray compatible' could cause it to trigger for pure NumPy/Dask/Xarray tasks better served by skills dedicated to those libraries.

4 / 5

Total

15

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (785 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.