CtrlK
BlogDocsLog inGet started
Tessl Logo

hypogenic

Plans and audits use of ChicagoHAI HypoGeniC/HypoRefine for LLM-assisted hypothesis generation from labeled text datasets. Use for the `hypogenic` package, its task configs, hypothesis banks, or HypoBench datasets—not for manual hypothesis formulation or scientific validation.

73

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers executable, well-structured guidance with a strongly sequenced, validation-rich workflow and clean one-level-deep progressive disclosure into real bundle files. Its only notable weakness is mild verbosity in a few informational sections.

DimensionReasoningScore

Conciseness

The body is dense and assumes Claude's competence (no introductory explanations of what a dataset or LLM is), with dated version/hash detail justified by reproducibility rather than padding; a few sections (privacy gate, upstream CLI facts) could be trimmed, so it sits above the 'mostly efficient' 3 anchor but below a maximally lean 5.

4 / 5

Actionability

It provides fully executable, copy-paste-ready commands for every bundled tool (validate_config, audit_dataset, plan_run, inspect_outputs, evaluate_local) with concrete flags and example paths covering the common cases, matching the 'fully executable' anchor.

5 / 5

Workflow Clarity

The eight-step 'Default workflow' is clearly sequenced with explicit validation checkpoints (dataset audit, checksum/split-leakage checks, cost/run plan, separate confirmation before external calls, local inspection) and feedback guidance, satisfying the batch/destructive validation requirement rather than the capped-3 case.

5 / 5

Progressive Disclosure

SKILL.md is a clear overview with well-signaled, one-level-deep references to verified bundle files (references/configuration.md, upstream.md, datasets.md, evaluation.md, security.md, sources.md and scripts/*.py), all of which exist, yielding easy navigation with no nested-reference indirection.

5 / 5

Total

19

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is concrete, trigger-rich, and answers both what and when with an explicit negative boundary that sharply distinguishes it from related skills. It is slightly short of maximal specificity and trigger-synonym coverage but otherwise strong.

DimensionReasoningScore

Specificity

Names the domain and several concrete capabilities ('Plans and audits use of ChicagoHAI HypoGeniC/HypoRefine for LLM-assisted hypothesis generation from labeled text datasets') plus concrete artifacts ('task configs, hypothesis banks, or HypoBench datasets'), with only minor coverage gaps, matching the 'several specific actions' anchor rather than the fully comprehensive 5.

4 / 5

Completeness

It explicitly answers both 'what' ('Plans and audits use of... hypothesis generation from labeled text datasets') and 'when' ('Use for the `hypogenic` package, its task configs, hypothesis banks, or HypoBench datasets') with concrete trigger phrases and an explicit negative boundary.

5 / 5

Trigger Term Quality

The 'Use for the `hypogenic` package, its task configs, hypothesis banks, or HypoBench datasets' clause supplies good natural trigger phrases users would say, but is missing some common synonyms/variations, sitting above the 3 anchor but below fully comprehensive 5 coverage.

4 / 5

Distinctiveness Conflict Risk

It carves a clear niche (the ChicagoHAI hypogenic package) and adds an explicit exclusion ('not for manual hypothesis formulation or scientific validation') that minimizes conflict with adjacent skills, matching the distinct-niche anchor.

5 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
K-Dense-AI/scientific-agent-skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.