CtrlK
BlogDocsLog inGet started
Tessl Logo

exploratory-data-analysis

Analyze scientific data files across 200+ formats at the depth the user requests. Detect file type, assess structure, quality, and statistics, and create reports or visualizations only when they are requested or materially needed. Covers chemistry, bioinformatics, microscopy, spectroscopy, proteomics, metabolomics, and general scientific data formats.

58

Quality

68%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/coding/exploratory-data-analysis/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with a genuinely progressive-disclosure design — a lean overview pointing to six large, real reference files, a runnable script, and a report template, all cited with accurate paths and grep-based lookup instructions. Weaknesses are moderate: redundant capability/best-practices sections and comment-style examples cost tokens without adding executable value, and validation checkpoints live in Best Practices rather than in the workflow itself.

Suggestions

Cut the 'Key Capabilities' bullet list and the 'Report Generation' best practices ('Be comprehensive / Be specific / Be actionable') — they restate the description and generic advice Claude already knows.

Convert at least one Example (e.g., the FASTQ one) from comment-style outline to executable code that computes and prints read count, GC content, and quality stats.

Fold the metadata validation check ('Cross-check metadata consistency (e.g., stated dimensions vs actual data)') into Step 3 of the workflow as an explicit checkpoint rather than leaving it in Best Practices.

DimensionReasoningScore

Conciseness

Mostly efficient — it assumes Claude's competence (never explains what FASTQ or pandas is) and gives library names without library tutorials — but several sections pad: the 'Key Capabilities' bullet list restates the description, 'Report Generation' best practices ('Be comprehensive', 'Be specific', 'Be actionable') are generic advice Claude already knows, and the three Examples are comment-style walkthroughs that largely repeat the workflow. It is noticeably tighter than a 2 but has enough unnecessary material to fall short of 4.

3 / 5

Actionability

Provides mostly executable guidance: the analyzer invocation 'python scripts/eda_analyzer.py <filepath> [output.md]', a concrete regex ('### \.pdb[^#]*?(?=###|\Z)') for reference lookup, an ImportError fallback with an install command, and per-datatype analysis checklists. Below 5 because the Examples are pseudocode outlines ('# Calculate: read count, length distribution...') rather than copy-paste runnable code, and the per-format analysis steps are named but not shown.

4 / 5

Workflow Clarity

The five-step workflow (detect → load reference → analyze → present → save-only-if-requested) is clearly sequenced with explicit decision points ('For a narrow request, return the result inline and stop'), plus a Troubleshooting section covering import errors, unknown extensions, and large files. It is a read-only analysis skill, so the destructive/batch validation cap does not apply; it misses 5 only because validation checkpoints (e.g., cross-checking stated vs actual dimensions) appear in Best Practices rather than being wired into the workflow as explicit checkpoints.

4 / 5

Progressive Disclosure

Scored against the actual bundle: six large reference files, one script, and one asset all exist at the paths cited, and the body signals each with its exact path and purpose ('Reference file: references/chemistry_molecular_formats.md', 'assets/report_template.md', 'scripts/eda_analyzer.py'), plus explicit instructions to grep sections rather than load whole files. References are one level deep (spot-check confirms a flat '### .extension' structure, no nested pointers), and the Resources section indexes every bundle file — matching the 'clear overview with well-signaled one-level-deep references' anchor.

5 / 5

Total

16

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is specific, third-person, and uncluttered, with a well-scoped multi-domain niche and solid trigger vocabulary. Its main structural weakness is the absence of any explicit 'Use when...' guidance, which both caps completeness and slightly weakens its usefulness for skill routing.

Suggestions

Add an explicit trigger clause, e.g., 'Use when the user provides or asks about a scientific data file, asks to explore, analyze, or summarize one, or wants a data quality assessment.'

Include a few representative file extensions or synonyms (e.g., .fastq, .mzML, .csv, 'EDA') so the description matches natural user phrasing and file-based triggers.

Trim the conditional phrasing about when reports are created ('only when they are requested or materially needed') — it is output policy better placed in the body, and it dilutes the description.

DimensionReasoningScore

Specificity

Quotes several concrete actions — 'Detect file type', 'assess structure, quality, and statistics', 'create reports or visualizations' — which matches the anchor 'lists several specific actions; minor gaps in coverage'. It falls short of a 5 because no file extensions or format-level specifics are named, and sits above 3 because more than 1-2 distinct actions are given.

4 / 5

Completeness

The 'what' is clearly stated ('Detect file type, assess structure, quality, and statistics, and create reports or visualizations'), but there is no 'Use when...' clause or equivalent explicit trigger guidance — 'only when they are requested or materially needed' constrains output, not when to invoke the skill. Per the rubric guideline, a missing 'Use when' clause caps completeness at 3; it is not a 2 because the 'what' half is clear and specific.

3 / 5

Trigger Term Quality

Good keyword coverage including natural terms users would say — 'analyze', 'scientific data files', 'reports', 'visualizations' — plus seven domain names (chemistry, bioinformatics, microscopy, spectroscopy, proteomics, metabolomics). A few natural terms are missing (e.g., 'explore', 'summarize', 'EDA', and file extensions like .csv or .fastq), so it is not the comprehensive 5.

4 / 5

Distinctiveness Conflict Risk

The scoped domain list ('Covers chemistry, bioinformatics, microscopy, spectroscopy, proteomics, metabolomics, and general scientific data formats') carves a clear scientific-data niche with distinct triggers. Minor overlap risk remains with generic data-analysis or spreadsheet skills since no extensions anchor the niche, keeping it below 5.

4 / 5

Total

15

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.