CtrlK
BlogDocsLog inGet started
Tessl Logo

toxicity-structure-alert

Analyze data with `toxicity-structure-alert` using a reproducible workflow, explicit validation, and structured outputs for review-ready interpretation.

41

Quality

52%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./scientific-skills/Data Analysis/toxicity-structure-alert/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

60%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The domain core is strong — an accurate alert-structure table, executable CLI/Python examples that match the packaged script, and a realistic output contract. The skill is dragged down by template boilerplate: roughly a third of the body is redundant governance prose with dangling self-references, and a few commands are misleading (non-SMILES audit input, nonexistent requirements.txt and hardcoded cd path).

Suggestions

Collapse the overlapping governance sections (Workflow, Output Requirements, Output Contract, Response Template, Inputs to Collect, Input Validation, Error Handling, Validation and Safety Rules) into one concise workflow + error-handling section, and delete the 'See `## X` above' filler lines.

Fix the misleading commands: replace the non-SMILES 'Audit validation sample ...' input with a real SMILES string, remove the hardcoded `cd "20260318/scientific-skills/..."` path, and either add a requirements.txt or state dependencies inline (`pip install rdkit`).

Name and link the reference file explicitly (e.g. 'See [references/runtime_checklist.md](references/runtime_checklist.md)') instead of only pointing at the `references/` directory.

DimensionReasoningScore

Conciseness

Beyond the genuinely useful core (alert table, CLI usage, output format), the body carries heavy template padding: three dangling 'See `## Features`/`## Usage`/`## Workflow` above' pointers, a 'Key Features' section that repeats the frontmatter description verbatim, and ~8 overlapping governance sections (Workflow, Output Requirements, Output Contract, Response Template, Inputs to Collect, Input Validation, Error Handling, Validation and Safety Rules) restating the same rules, plus date-stamped lifecycle filler ('Next Review Date: 2026-03-06'). This matches 'Noticeably verbose; several unnecessary explanations or padded sections'. Not 3 because the padding is extensive rather than minor; not 1 because the domain reference material is concrete and does not explain concepts Claude already knows.

2 / 5

Actionability

The usage section is copy-paste ready and verified against the script: `python scripts/main.py -i "O=[N+]([O-])c1ccccc1"`, `-f json`, `-d full`, documented `--input/--format/--detail` flags, a working `ToxicityAlertScanner` Python API example, and a realistic JSON output sample. Minor gaps: the 'Audit-Ready' command passes an English clinical sentence as `--input` (not a valid SMILES), Prerequisites references a nonexistent `requirements.txt`, and Example Usage hardcodes a bundle path (`cd "20260318/scientific-skills/..."`) that does not exist. This fits 'Mostly executable guidance; concrete code or commands with minor gaps'. Not 5 because of those misleading/broken commands; not 3 because the primary examples are fully executable and cover common cases.

4 / 5

Workflow Clarity

The 5-step Workflow sequences confirmation, scope validation, execution, structured return, and an explicit fallback ('If execution fails ... switch to the fallback path and state exactly what blocked full completion'), with a non-destructive smoke check (`python -m py_compile scripts/main.py`) and Error Handling rules forming a validate-and-recover loop — matching 'Clear sequence with most checkpoints present; minor validation gaps'. Not 5 because the process is scattered across many redundant sections and the 'Example run plan' steps are abstract ('confirm ... edit ... run ... review'); not 3 because validation checkpoints and error-recovery feedback loops are explicitly present.

4 / 5

Progressive Disclosure

The bundle matches the body: `scripts/main.py` is the stated primary implementation surface and `references/` exists (runtime_checklist.md) with the body pointing to it generically ('Reference guidance: `references/` contains supporting rules, prompts, or checklists'), one level deep. Good structure overall with section headers, matching 'Good structure; most content is appropriately placed; references mostly clear; minor organization gaps'. Not 5 because the reference file is never named or linked directly (only the directory), and the self-referential 'See ## X above' pointers add navigation noise.

4 / 5

Total

14

/

20

Passed

Description

25%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is template-generated process language that omits the skill's actual purpose (scanning drug molecules/SMILES for known toxic structural alerts) and any 'use when' trigger guidance. It would rarely fire correctly and could collide with any other data-analysis skill.

Suggestions

State the concrete capability: e.g. 'Scan drug molecules (SMILES strings) for known toxic structural alerts (aromatic nitro, epoxide, hydrazine, ...) and report risk levels and recommendations.'

Add an explicit trigger clause with natural user vocabulary: 'Use when the user asks about toxicity screening, structural/toxicophore alerts, mutagenicity or carcinogenicity risk of a molecule, or provides a SMILES string to assess.'

Remove the generic process buzzwords ('reproducible workflow', 'review-ready interpretation') that apply to any analysis skill and dilute distinctiveness.

DimensionReasoningScore

Specificity

The only action named is the generic "Analyze data"; the skill's actual capability — identifying toxic structural alerts in drug molecules by scanning SMILES — is absent. The rest is process padding ("reproducible workflow", "explicit validation", "structured outputs for review-ready interpretation") with no concrete domain actions, matching the anchor 'Names the domain but actions are minimal or generic' — the domain is only hinted via the tool name. Not 3 because no concrete domain action (scan, identify, classify alerts) is stated; not 1 because the tool name and output characteristics do narrow it beyond pure abstraction.

2 / 5

Completeness

The 'what' is vague ("Analyze data" — what data, to what end?) and there is no 'when' clause or equivalent trigger guidance anywhere, which per the judging guidelines caps completeness at 3 and here matches anchor 2: 'Has a vague what and no when'. Not 3 because even the 'what' lacks a concrete capability statement; not 1 because it does state some output characteristics (validation, structured outputs).

2 / 5

Trigger Term Quality

Natural keywords a user would say ("structural alerts", "toxicity screening", "SMILES", "mutagenicity", "drug molecules") are missing; the only relevant terms ("toxicity", "structure", "alert") appear solely inside the backticked tool identifier `toxicity-structure-alert`, and "Analyze data" plus "validation"/"structured outputs" are generic. This fits 'One or two generic keywords; missing the natural phrases users say'. Not 3 because there is no independent keyword coverage beyond the tool name — no synonyms or natural phrasing.

2 / 5

Distinctiveness Conflict Risk

"Analyze data ... using a reproducible workflow, explicit validation, and structured outputs" would equally describe virtually any data-analysis skill in a bundle, creating high overlap/conflict risk; the sole distinguishing element is the tool name itself, matching 'Very broad; high overlap risk with many similar skills'. Not 3 because nothing beyond the tool identifier differentiates it from sibling analysis skills; not 1 because the tool name does anchor it to a toxicity niche.

2 / 5

Total

8

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 4 missing

Warning

Total

14

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.