CtrlK
BlogDocsLog inGet started
Tessl Logo

arboreto

Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.

65

Quality

80%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/coding/arboreto/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

67%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, actionable skill body with a genuine bundle (three reference files and a runnable script), clear troubleshooting, and reproducibility guidance. Main weaknesses are duplicated content between SKILL.md and the reference files, an undefined `analyze_consensus` call in an example, and the absence of explicit output-validation checkpoints.

Suggestions

Remove the duplicated install command and trim the Algorithm Selection and Distributed Computing sections to one short snippet each, pointing to `references/algorithms.md` and `references/distributed_computing.md` for the rest.

Replace the undefined `analyze_consensus(networks)` call with concrete consensus-filtering code (e.g., merge on (TF, target) and keep links appearing in a majority of seeds).

Add a short validation step after inference, such as checking the output DataFrame is non-empty and has the expected TF/target/importance columns before saving.

DimensionReasoningScore

Conciseness

Mostly efficient, but there is recurring avoidable duplication rather than merely minor trimmings: the install command appears twice (Quick Start and Installation), the Algorithm Selection section inlines a GRNBoost2-vs-GENIE3 comparison that duplicates `references/algorithms.md`, and the Distributed Computing section inlines three code blocks (default, custom local client, cluster) that overlap `references/distributed_computing.md`. This fits the 3 anchor ('mostly efficient but includes some unnecessary explanation or could be tightened') better than 4, where over-explanation would be only minor.

3 / 5

Actionability

Nearly all code is concrete, copy-paste ready (quick start, TF filtering, multi-condition loops) and there is a ready-to-run script with an exact CLI invocation: `python scripts/basic_grn_inference.py expression_data.tsv output_network.tsv --tf-file tfs.txt --seed 777`. It falls short of 5 because the Reproducibility example calls an undefined `analyze_consensus(networks)` that would raise NameError, and the pySCENIC integration snippet ends with 'See pySCENIC documentation' rather than a concrete next step.

4 / 5

Workflow Clarity

The single-task workflow (install → load data → infer → save → filter/interpret) is unambiguous via the Quick Start and use-case sections, and the Troubleshooting section supplies error-recovery guidance for memory, performance, Dask, and empty results ('Check data format (genes as columns), verify TF names match gene names'). It does not reach 5 because there are no explicit validation checkpoints (e.g., verify the output has expected columns/non-empty before saving), placing it at the 'clear sequence with most checkpoints; minor validation gaps' anchor.

4 / 5

Progressive Disclosure

Good structure: SKILL.md is an overview with well-signaled, real, one-level-deep references (`references/basic_inference.md`, `references/algorithms.md`, `references/distributed_computing.md` — all exist) plus a bundled script, and the one cross-link found in the references (basic_inference.md → algorithms.md) is a sibling link, not nested content. It falls short of 5 because substantial content that belongs in the references (algorithm comparison details, three distributed-computing code blocks) is inlined in SKILL.md alongside the pointers, a minor organization gap matching the 4 anchor.

4 / 5

Total

15

/

20

Passed

Description

92%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete capabilities with named algorithms, an explicit 'Use when' trigger covering both bulk and single-cell RNA-seq, and a clearly distinct niche. The only improvement space is adding a few ecosystem synonyms (e.g., SCENIC/pySCENIC) users might mention when they need this skill.

Suggestions

Add 'SCENIC' or 'pySCENIC' as a trigger synonym, since users often frame GRN inference requests in terms of the SCENIC pipeline.

Optionally mention 'co-expression' or 'network inference from expression' phrasing to catch users who don't use the term GRN.

DimensionReasoningScore

Specificity

Multiple concrete actions are named with specifics: 'Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3)', 'identify transcription factor-target gene relationships', and 'Supports distributed computation for large-scale datasets' — comprehensive coverage of the library's capability surface. Score 4 would apply only if there were minor gaps in coverage, but the named algorithms, data types, and scale behavior leave no significant gap.

5 / 5

Completeness

Explicitly answers both questions: what — 'Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3)', and when — 'Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships'. This matches the 5 anchor with a concrete 'Use when...' trigger clause naming specific data modalities, so it is not the 4 anchor where 'when' could be more explicit.

5 / 5

Trigger Term Quality

Good natural-keyword coverage: 'gene regulatory networks', 'GRNs', 'gene expression data', 'bulk RNA-seq', 'single-cell RNA-seq', 'transcriptomics', plus algorithm names GRNBoost2/GENIE3 that practitioners would say. A few natural terms are missing — most notably 'SCENIC'/'pySCENIC' (the dominant ecosystem keyword for this task) and 'co-expression' — so it does not reach the comprehensive-synonyms anchor of 5, while it clearly exceeds the 'some relevant keywords but missing variations' anchor of 3.

4 / 5

Distinctiveness Conflict Risk

A clear niche (GRN inference from transcriptomics data) with distinct triggers — algorithm names, RNA-seq modalities, TF-target relationships — that would not naturally collide with other skills. It is not the 4 anchor because there is no meaningful overlap with a closely related skill's trigger surface; generic-bioinformatics conflict risk is minimal.

5 / 5

Total

19

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.