CtrlK
BlogDocsLog inGet started
Tessl Logo

scikit-survival

A comprehensive toolkit for survival analysis and time-to-event modeling in Python using scikit-survival; use it when you need to model censored time-to-event outcomes, fit Cox/RSF/GB models or Survival SVMs, evaluate with C-index/Brier score, or handle competing risks.

65

Quality

82%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

71%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-built body with an excellent runnable end-to-end example and concrete API-level guidance throughout. The main weaknesses are the hedged, unlinked reference listing ('may exist under') with inline duplication of reference-file topics, and minor actionability gaps in the competing-risks and hyperparameter-tuning sections.

Suggestions

Replace the hedged blockquote with confident, linked, per-topic pointers placed where each topic arises — e.g., under the Cox heuristics: 'Penalized Cox details: see [references/cox-models.md](references/cox-models.md)' — and remove the 'may exist' phrasing.

Drop or fill the empty param_grid placeholder block; either give a working tuning example (e.g., tuning GradientBoostingSurvivalAnalysis learning_rate) or cut the branch entirely.

Make the competing-risks snippet concrete by showing how event types are actually encoded (e.g., a structured array with an integer event field and a real cumulative-incidence call) instead of '# y must encode event types appropriately'.

DimensionReasoningScore

Conciseness

The body is efficient — it lists model names, API calls, and short heuristics without explaining concepts Claude already knows — but includes padding that could be trimmed: an empty 'param_grid' block with placeholder comments ('Example placeholder; remove if unsupported in your installed version'), and overlap between 'When to Use', 'Key Features', and 'Implementation Details'. This fits anchor 4 (efficient, minor instances that could be trimmed) rather than anchor 5 (every token earns its place).

4 / 5

Actionability

The Example Usage is fully executable with a real dataset ('X, y = load_breast_cancer()'), a pipeline, and metric calls, and Implementation Details give concrete Surv construction and metric snippets. Minor gaps remain: the competing-risks snippet is vague ('# y must encode event types appropriately for competing risks workflows' with no concrete construction pattern), and the tuning section defers with 'If your version exposes regularization parameters, tune them here'. Mostly executable with minor gaps — anchor 4, not the copy-paste-ready common-case coverage of anchor 5.

4 / 5

Workflow Clarity

The main example is a clearly numbered sequence (load → split → pipeline → tune → predict → evaluate) and preprocessing guidance includes checks ('ensure non-negative times; verify enough events relative to feature count'). However, there are no explicit validation checkpoints or error-recovery loops in the evaluation workflow — e.g., what to check before interpreting IPCW metrics. Clear sequence with most checkpoints present but minor validation gaps — anchor 4; this is an analysis skill, so the destructive/batch cap at 3 does not apply.

4 / 5

Progressive Disclosure

Six reference files exist and are listed, but they are signaled weakly: a blockquote saying guides 'may exist under:' (hedged, not confident navigation), rendered as plain paths rather than links, and not tied to the body topics that duplicate their coverage (inline model-selection heuristics, metric snippets, and competing-risks code that also live in the reference files). References present but not clearly signaled with content that arguably should be separate inline — anchor 3, below anchor 4's 'references mostly clear'.

3 / 5

Total

15

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that explicitly answers both what the skill does and when to use it, with concrete capability phrases and a distinct niche. Its only weaknesses are the second-person phrasing (penalized under specificity) and missing natural synonyms like Kaplan-Meier in the trigger terms.

Suggestions

Rewrite in third person to avoid the second-person penalty: 'A comprehensive toolkit for survival analysis and time-to-event modeling in Python using scikit-survival. Use when modeling censored time-to-event outcomes...' instead of 'use it when you need to...'.

Add widely-used natural trigger terms such as 'Kaplan-Meier' and 'hazard ratios' to broaden the keyword coverage to anchor-5 level.

DimensionReasoningScore

Specificity

Concrete actions are listed comprehensively — 'model censored time-to-event outcomes, fit Cox/RSF/GB models or Survival SVMs, evaluate with C-index/Brier score, or handle competing risks' — matching the anchor-5 example. However, the rubric penalizes second-person voice, and the description slips into it ('use it when you need to model...'), reducing this from 5 to 4.

4 / 5

Completeness

Both 'what' and 'when' are explicitly answered: 'A comprehensive toolkit for survival analysis and time-to-event modeling in Python using scikit-survival; use it when you need to...' with concrete trigger phrases (censored outcomes, Cox/RSF/GB models, C-index/Brier score, competing risks). This matches the anchor-5 example structure exactly.

5 / 5

Trigger Term Quality

Good coverage of natural terms users would say — 'survival analysis', 'time-to-event', 'censored', 'Cox', 'C-index', 'Brier score', 'competing risks' — but common synonyms a user might naturally mention are missing (e.g., 'Kaplan-Meier', 'hazard ratio', 'time-to-event data'), which fits anchor 4 rather than the comprehensive synonym coverage of anchor 5.

4 / 5

Distinctiveness Conflict Risk

Clear niche (survival analysis with scikit-survival) with distinct triggers — censored time-to-event data, Cox models, Survival SVMs, C-index — that would not plausibly fire for unrelated skills. Minimal conflict risk, matching anchor 5.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

15

/

16

Passed

Repository
aipoch/medical-research-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.