CtrlK
BlogDocsLog inGet started
Tessl Logo

curated-bio-datasets

Guide to accessing curated biological datasets for computational biology. COSMIC cancer data, GTEx expression, GWAS catalog, GeneBass exome variants, BioGRID interactions, MSigDB gene sets, DisGeNET disease-gene associations, and GO ontology. For specific database APIs use individual database skills (cosmic-database, gwas-database, etc.).

51

Quality

57%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/biology/curated-bio-datasets/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

53%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-sectioned and highly actionable, with real code, column names, and endpoints for all eight databases. Its core structural flaw is that it is a monolithic ~560-line inline implementation while three equivalent scripts sit unreferenced in scripts/, duplicating functionality and wasting the context window; workflows also lack validation checkpoints.

Suggestions

Replace the inline COSMIC, BioGRID/PPI, and MSigDB code sections with brief overviews plus explicit references to the existing bundle scripts (e.g. 'Run: python scripts/build_ppi_network.py --biogrid-file ... --gene-list ...'), making the provided scripts discoverable.

De-duplicate 'parse_gmt' (defined in both Quick Start and section 6 with different return formats) and trim print-boilerplate from the remaining functions.

Turn the 'Typical Workflows' into sequenced steps with validation checkpoints (e.g. download → verify file/row count → load → filter → report) so failures are caught early.

DimensionReasoningScore

Conciseness

Prose is lean and code-dense rather than padded, but the ~560-line body could be tightened: 'parse_gmt' is defined twice (Quick Start and section 6) with divergent return shapes, and every function carries verbose print-based reporting boilerplate. Matches 'mostly efficient but could be tightened' rather than the minor-trim level of 4.

3 / 5

Actionability

Each section ships concrete, executable pandas/networkx/requests code with real column names ('Tumour Types(Somatic)', 'MAPPED_GENE'), download URLs, and file names — close to copy-paste ready. It stays at 4 rather than 5 because acquiring the data is only described in comments (e.g. 'requires free registration') with no runnable download step, and code assumes the local files already exist.

4 / 5

Workflow Clarity

The three 'Typical Workflows' are 2–3-line snippets, not sequenced processes; there are no validation checkpoints (e.g. verify a download parsed to expected row counts) or explicit feedback loops. The Troubleshooting section provides some problem/solution recovery, which keeps this at 3 ('sequence present but checkpoints missing') rather than 2.

3 / 5

Progressive Disclosure

The bundle provides three substantive scripts (scripts/build_ppi_network.py, scripts/download_cosmic.py, scripts/parse_msigdb.py) that duplicate the PPI, COSMIC, and MSigDB sections, yet the body never references any of them and instead inlines ~500 lines of code that clearly belongs in those files. Per the rubric guideline on scoring against the actual bundle structure, the provided files are orphaned and undiscoverable — 'content that clearly belongs in separate files is inlined' at anchor 2, not 3, since no references to the bundle exist at all.

2 / 5

Total

12

/

20

Passed

Description

61%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description names a strong, concrete set of eight databases that serve as effective trigger terms and clearly delineates its boundary with dedicated database skills. Its main weaknesses are the absence of an explicit 'Use when...' trigger clause and reliance on a single generic action verb rather than itemized capabilities.

Suggestions

Add an explicit trigger clause, e.g. 'Use when downloading or parsing COSMIC, GTEx, GWAS Catalog, BioGRID, MSigDB, or DisGeNET data, or when building PPI networks or running pathway enrichment.'

Replace the generic verb 'accessing' with specific actions per dataset (e.g. 'parse COSMIC census CSVs, query the GTEx API, load MSigDB GMT files, build BioGRID PPI networks').

Include natural synonyms and file formats users mention (GMT, TPM, eQTL, .gmt, pathway enrichment) to round out trigger coverage.

DimensionReasoningScore

Specificity

The description names eight concrete datasets ('COSMIC cancer data, GTEx expression, GWAS catalog, GeneBass exome variants, BioGRID interactions, MSigDB gene sets, DisGeNET disease-gene associations, and GO ontology') but relies on a single generic action verb 'accessing' — it names objects comprehensively, not actions. This matches 'Names domain and 1-2 concrete actions' better than 'Lists several specific actions', and it is clearly above 'Names the domain but actions are minimal'.

3 / 5

Completeness

The 'what' is clear ('Guide to accessing curated biological datasets...') but there is no 'Use when...' clause or equivalent positive trigger guidance; the final sentence ('For specific database APIs use individual database skills') is a delegation boundary, not a when-to-use trigger. Per the rubric guideline, a missing trigger clause caps completeness at 3.

3 / 5

Trigger Term Quality

Database names like 'COSMIC', 'GTEx', 'GWAS catalog', 'BioGRID', 'MSigDB', and 'DisGeNET' are exactly the natural terms users in computational biology would say. Falls short of a 5 because common synonyms and file-format triggers (GMT, TPM, pathway enrichment, eQTL, .gmt) are absent.

4 / 5

Distinctiveness Conflict Risk

It carves out a distinct niche (bulk curated dataset access) and explicitly delegates API-level work to sibling skills ('cosmic-database, gwas-database, etc.'). Minor overlap remains: a user asking about COSMIC or GWAS data could plausibly trigger either this skill or the dedicated one, so it does not reach a 5.

4 / 5

Total

14

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (570 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.