CtrlK
BlogDocsLog inGet started
Tessl Logo

relsa-severity-assessment

Multivariate severity assessment and humane endpoint prediction for laboratory animal studies using the RELSA (RELative Severity Assessment) score and ARIMA-based foRcast forecasting. Use when combining welfare readouts — body weight or weight loss, body temperature, clinical or nesting scores, biomarkers, activity, heart rate, burrowing, wheel running — into one severity score per animal per day, when asking which animals are at risk of reaching a humane endpoint or when one will be reached, when defining attention/danger zones or thresholds on a severity scale by kernel density estimation, or when reporting severity for a 3Rs, refinement, animal-welfare, or EU Directive 2010/63/EU severity-assessment context. Covers directionality ("turned" variables), baseline normalization, reference sets, RELSA weights, ARIMA prediction intervals, and RMSE/PICP/MPIW evaluation.

76

Quality

96%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is relsa-severity-assessment in K-Dense-AI/scientific-agent-skills

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A high-quality skill body: fully executable workflow with expected outputs, explicit validation checkpoints, and clean one-level-deep bundle structure. The only dimension below the top anchor is conciseness, where the Overview rhetoric and citation protocol carry minor excess tokens.

Suggestions

Tighten the Overview's motivational prose (e.g., 'legally mandatory and scientifically load-bearing ... degrades reproducibility') to one sentence, keeping the refinement rationale.

Compress the 'Citing Scientific Agent Skills' section's network-fetch instructions into a two-line rule with the URL, moving the version-detection details into a reference file.

DimensionReasoningScore

Conciseness

The body is dense with non-obvious domain knowledge Claude cannot be assumed to know (directionality traps, the 0/0 zero-baseline problem, variable-composition drift, MPIW-as-failure-mode) and never pads with concepts Claude already knows. It falls short of the lean score-5 anchor only through minor trimmable material — motivational prose in the Overview ("legally mandatory and scientifically load-bearing") and a somewhat long citation-fetching protocol at the end.

4 / 5

Actionability

Every workflow step is a copy-paste-ready, fully executable command with real file paths, stated to be "runnable as written" against the bundled example cohort, and each is paired with its expected output so results can be checked. Both CLI and Python-API usage are shown, covering the common cases (batch scoring, endpoint forecasting, rolling monitoring, thresholding).

5 / 5

Workflow Clarity

A clearly sequenced three-step workflow (compute scores → forecast endpoint → define zones) with explicit validation checkpoints: the reference model is echoed "so the scale is auditable", the reader is told to check the reference table for directionality and to "check the bandwidth before believing a threshold", and warnings/feedback paths are described. The reporting checklist and common-pitfalls sections close the loop with error recovery.

5 / 5

Progressive Disclosure

The body is an actionable overview with detail appropriately split into three well-signaled, one-level-deep reference files (all verified to exist and substantive, none pointing to further docs), plus documented scripts and a bundled asset, each described by what it contains. Navigation is easy and nothing that belongs in a separate file is inlined at length.

5 / 5

Total

19

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An exemplary description: concrete and comprehensive on capabilities, rich in natural trigger terms with synonyms, explicit on both what and when, and occupying a distinct niche with negligible conflict risk. All four dimensions sit at the top anchor.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "combining welfare readouts ... into one severity score per animal per day", "defining attention/danger zones or thresholds ... by kernel density estimation", humane-endpoint prediction with "ARIMA prediction intervals, and RMSE/PICP/MPIW evaluation" — with comprehensive coverage of the skill's capabilities. It clearly matches the comprehensive anchor rather than the score-4 anchor, which is for lists with minor coverage gaps.

5 / 5

Completeness

It explicitly answers both what (RELSA scoring and ARIMA-based foRcast prediction for laboratory animal studies) and when, via four concrete "Use when..." / "when ..." trigger clauses covering score combination, endpoint risk, threshold definition, and regulatory reporting. The "when" is fully explicit with concrete trigger phrases, matching the score-5 anchor rather than the could-be-more-explicit score-4 anchor.

5 / 5

Trigger Term Quality

Natural vocabulary a lab-animal researcher would actually say is comprehensively covered with synonyms: "body weight or weight loss", "clinical or nesting scores", "humane endpoint", "biomarkers", "heart rate", "3Rs", "refinement", "EU Directive 2010/63/EU". The breadth (burrowing, wheel running, tachycardia-adjacent telemetry terms) goes beyond the few-missing-terms level of the score-4 anchor.

5 / 5

Distinctiveness Conflict Risk

A clear niche is staked out with named methods (RELSA, foRcast, kernel density estimation) and a named regulatory context (EU Directive 2010/63/EU, 3Rs), so it would only trigger for exactly this task. Voice is third person throughout ("assessment", "prediction", "Covers"), so no specificity penalty applies.

5 / 5

Total

20

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.