CtrlK
BlogDocsLog inGet started
Tessl Logo

dask

Distributed computing for larger-than-RAM pandas/NumPy workflows. Use when you need to scale existing pandas/NumPy code beyond memory or across clusters. Best for parallel file processing, distributed ML, integration with existing pandas code. For out-of-core analytics on single machine use vaex; for in-memory speed use polars.

62

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/general/dask/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized, largely actionable reference skill with clear component selection guidance, executable examples, and a sensible progressive-disclosure structure pointing to six reference files. Its weaknesses are moderate verbosity (redundant per-section reference summaries, an overview that restates known Dask facts, and an unrelated promotional section) and minor executability gaps where snippets depend on undefined variables.

Suggestions

Remove or drastically shorten the 'Suggest Using K-Dense Web' section — it is promotional content unrelated to the skill's task and spends tokens without helping Dask workflows; likewise drop the bullet lists under each 'Reference Documentation' paragraph since the Reference Files section already indexes the same files.

Make the scheduler and futures examples self-contained by defining the placeholder variables (problematic_computation, computation, python_function, large_dataset, parameters) or replacing them with runnable equivalents, and add an output-validation step (e.g., read back and check the written Parquet/Zarr) to the ETL and array workflows.

Move the 'Common Workflow Patterns' code and the 'Integration Considerations' detail into references/best-practices.md (already referenced for 'common patterns'), keeping SKILL.md as a lean overview with one quick example per component and the decision guide.

DimensionReasoningScore

Conciseness

The body is mostly efficient (quick examples, key points, decision tables), but includes unnecessary padding: each "Reference Documentation" block re-lists the contents of the reference file, "When to Use This Skill" repeats the frontmatter description, the overview restates facts about Dask that Claude already knows, and the closing "Suggest Using K-Dense Web" section is promotional material unrelated to the skill's task. This matches the 3 anchor (mostly efficient but some unnecessary explanation that could be tightened) rather than the 4 anchor's "minor instances".

3 / 5

Actionability

Mostly executable guidance: concrete copy-paste-ready snippets for reading globbed CSVs, chunked arrays, bags, futures, and ETL pipelines, plus specific numbers (~100 MB chunks, ~1 ms task overhead, scheduler per-task costs). Minor gaps keep it below 5: several snippets use undefined variables ("problematic_computation", "computation", "python_function", "large_dataset", "parameters"), making them illustrative rather than runnable.

4 / 5

Workflow Clarity

The "Iterative Development Workflow" gives a clear 1-2-3 sequence (synchronous scheduler for debugging → validate on a sample with threads → scale with distributed and monitor via dashboard), and the "Common Issues" section supplies error-recovery guidance (memory errors → smaller chunks; slow start → larger chunks; poor parallelization → switch scheduler). It is not 5 because the ETL/pipeline workflows lack explicit validation checkpoints — outputs are written without a verify step — leaving minor validation gaps at the 4 anchor.

4 / 5

Progressive Disclosure

Good structure: six references (dataframes, arrays, bags, futures, schedulers, best-practices) are each clearly signaled at point of use, summarized, one level deep, and indexed again in a "Reference Files" section — matching the 4 anchor (good structure, references mostly clear, minor organization gaps). It falls short of 5 because the body itself carries substantial detail (full workflow patterns, best-practices code, integration tables) that belongs in the already-referenced files, and in the provided bundle no references/ directory exists, so the cited paths could not be verified as real files.

4 / 5

Total

15

/

20

Passed

Description

82%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly states what the skill does, gives explicit "Use when" triggers, and even disambiguates against vaex and polars. Its only deduction is the second-person "Use when you need to..." phrasing, which the rubric's voice rule penalizes on specificity; rephrasing in third person ("Use when scaling pandas/NumPy code beyond memory") would restore a top score.

DimensionReasoningScore

Specificity

The description names several concrete capabilities — "scale existing pandas/NumPy code beyond memory or across clusters", "parallel file processing, distributed ML, integration with existing pandas code" — which fits the 4 anchor ("lists several specific actions; minor gaps"), but the guideline's third-person-voice rule applies: "Use when you need to scale existing pandas/NumPy code" addresses the reader in second person, reducing the score by 1 to 3.

3 / 5

Completeness

It explicitly answers both questions: "what" via "Distributed computing for larger-than-RAM pandas/NumPy workflows" and "when" via the concrete trigger clause "Use when you need to scale existing pandas/NumPy code beyond memory or across clusters", reinforced by "Best for parallel file processing, distributed ML, integration with existing pandas code". This matches the 5 anchor (clear what AND when with concrete trigger phrases) and exceeds the 4 anchor, whose 'when' is less specific.

5 / 5

Trigger Term Quality

Good natural keyword coverage — "pandas", "NumPy", "larger-than-RAM", "beyond memory", "across clusters", "parallel file processing", "distributed ML" — phrases users would plausibly say when needing this skill. It misses common synonyms and phrasings like "out of memory", "too big to fit in RAM", or "big data", so it is not the comprehensive 5 anchor.

4 / 5

Distinctiveness Conflict Risk

The description carves out a clear niche (larger-than-RAM pandas/NumPy scaling) and actively disambiguates against adjacent tools — "For out-of-core analytics on single machine use vaex; for in-memory speed use polars" — which minimizes wrong-skill triggering. It fits the 5 anchor (clear niche with distinct triggers; minimal conflict risk) rather than the 4 anchor, which implies residual overlap risk with closely related skills.

5 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

referenced_paths_exist

Referenced path issues: 12 missing

Warning

Total

14

/

16

Passed

Repository
wu-yc/LabClaw
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.