CtrlK
BlogDocsLog inGet started
Tessl Logo

hypothesis-generation

Formulate evidence-bounded scientific questions, candidate hypotheses, rival explanations, causal or associational claims, discriminating predictions, measurements, and preregistration-ready analysis plans. Use when turning observations or preliminary findings into transparent, testable research plans without treating hypotheses as facts.

71

Quality

87%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

The canonical home for this skill is hypothesis-generation in K-Dense-AI/scientific-agent-skills

SKILL.md
Quality
Evals
Security

Quality

Content

86%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-structured, highly actionable workflow skill: clear 12-step sequence with safety gates, real validation scripts, and exemplary one-level-deep progressive disclosure backed by a complete bundle. Main weaknesses are mild verbosity from explaining known concepts and implicit rather than explicit error-recovery loops.

Suggestions

Trim explanations of concepts Claude already knows (FINER mnemonic, Platt's strong inference, reproducibility vs. replicability definitions), keeping only the skill-specific guidance on how to apply them.

Add an explicit fix-and-re-validate loop after the bundled validation commands (e.g., 'If exit code is 1, review the reported errors, fix the record, and re-run') to complete the feedback cycle.

Consider moving the citation/versioning instructions at the end into a reference file to keep the body focused on the workflow.

DimensionReasoningScore

Conciseness

The body is dense and mostly skill-specific procedural guidance (boundaries, workflow steps, tool index) with little padding, but stretches explaining concepts Claude already knows — the FINER mnemonic expansion, Platt's strong inference, and the reproducibility-vs-replicability distinction — could be trimmed. Not 3 because the bulk of the content is non-obvious procedural instruction; not 5 because these explanatory asides do consume tokens without adding skill-specific value.

4 / 5

Actionability

Guidance is fully executable: concrete copy-paste commands ("python3 scripts/check_operationalization.py local-operationalization.json"), a tool index table mapping tasks to real bundled assets and scripts, explicit exit-code semantics (0/1/2), and per-step record/checklist requirements. As an instruction-only skill its guidance is concrete and copy-paste ready.

5 / 5

Workflow Clarity

The 12-step workflow is clearly sequenced, starts with an explicit scope/safety gate, includes per-step checklist requirements, and provides validation via bundled scripts with documented exit codes. Not 5 because error-recovery loops after a failed validation (fix and re-run) are only implied by exit codes rather than spelled out.

4 / 5

Progressive Disclosure

The body is an overview with detail pushed to ten annotated, one-level-deep reference files plus a tool index; every referenced path (assets/*, references/*, scripts/*) resolves to a real file in the bundle, and references are clearly signaled with descriptions of what each contains. Easy to navigate with no nesting.

5 / 5

Total

18

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with concrete, comprehensive action coverage and an explicit 'Use when' trigger in proper third-person voice. Trigger-term coverage is good but could add common user synonyms, and it is mostly though not perfectly distinct from adjacent research-design skills.

Suggestions

Add a few natural user synonyms or phrasings (e.g., 'study design', 'experiment planning', 'research questions') to broaden trigger-term coverage.

Sharpen the 'when' clause to explicitly name the artifacts users ask for (e.g., 'Use when asked to generate hypotheses, rival explanations, or a preregistration plan') to further reduce overlap with general research-design or statistics skills.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "Formulate evidence-bounded scientific questions, candidate hypotheses, rival explanations, causal or associational claims, discriminating predictions, measurements, and preregistration-ready analysis plans" — with comprehensive coverage and no vague filler. Not 4 because the action list is both broad and fully concrete rather than having minor gaps.

5 / 5

Completeness

It explicitly answers both what ("Formulate evidence-bounded scientific questions...") and when ("Use when turning observations or preliminary findings into transparent, testable research plans") with a concrete trigger phrase in third-person voice. Clearly matches the anchor-5 example structure.

5 / 5

Trigger Term Quality

Natural user terms like "hypotheses", "scientific questions", "research plans", "observations", and "testable" are present, but common phrasings such as "experiment design", "study design", or "research questions" are underrepresented. Not 3 because keyword coverage is good; not 5 because synonym and variation coverage is incomplete.

4 / 5

Distinctiveness Conflict Risk

The scientific hypothesis-generation niche is clear with distinct triggers, but terms like "measurements" and "preregistration-ready analysis plans" carry minor overlap risk with general research-design or statistics skills. Not 5 because overlap with adjacent design/stats skills is possible but small.

4 / 5

Total

18

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
K-Dense-AI/claude-scientific-writer
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.