CtrlK
BlogDocsLog inGet started
Tessl Logo

idea-evaluator

Evaluates a preliminary research idea against a five-dimension framework (Higher, Faster, Stronger, Cheaper, Broader) plus idea-lifecycle and student-capability matching, paradigm-shift probing, and a fatal-flaws audit. Returns a reviewer-style verdict; non-STEM ideas route to substitute frameworks. Use when the user has a draft research idea and asks whether it is worth pursuing, asks to 'evaluate this idea', 'score this idea', 'assess feasibility', 'novelty check', 'is this a good research direction', or before committing to a paper scope.

72

Quality

91%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The content is a well-engineered procedural skill with standout workflow clarity (sequenced steps, early gate, integrity-gate checklist, feedback loop) and solid actionability via a complete output template. Its main weakness is progressive disclosure: over half the bundle's reference files are orphaned and unreachable through the skill's own navigation.

Suggestions

Link the orphaned per-principle files from references/paradigm-shift-probe.md (e.g., under each of the four probing principles, add 'See paradigm-first-principles.md / paradigm-elephant.md / paradigm-technology-cycle.md / paradigm-hamming.md for deeper treatment') so the deep dives become navigable.

Reference references/worked-examples.md from the Step 4 scoring section (or from five-dimensions.md) so the example evaluations are reachable, and reference paradigm-examples.md from the paradigm-shift probe step.

Tighten the Overview paragraph by dropping the motivational framing ('The goal is to kill weak ideas before the student invests months...') and keep only the one-sentence positioning, to lift conciseness toward the top anchor.

DimensionReasoningScore

Conciseness

The body is mostly efficient procedural instruction that assumes Claude's competence, but the Overview paragraph ('The goal is to kill weak ideas before the student invests months...') and a few rationale sentences are motivational framing that could be trimmed; not 5 because not every token strictly earns its place, not 3 because the bulk is purposeful rather than noticeably padded.

4 / 5

Actionability

Concrete, executable guidance throughout — an 8-step procedure, explicit scoring rules ('start every dimension at 5 and justify movement'), verdict thresholds, and a complete output-format template — but worked examples and several detection rubrics are deferred to reference files rather than surfaced inline, leaving minor gaps for the common cases.

4 / 5

Workflow Clarity

A clearly sequenced 8-step procedure with an explicit early gate (fatal-flaws audit before scoring), a short-circuit stop condition, and a 7-item Integrity-gate checklist tagged by enforceability class with a validate→downgrade feedback loop ('If any [inspection] check fails, downgrade the verdict and mark the affected output section'), matching the anchor-5 example of checkpoints, feedback loops, and checklists.

5 / 5

Progressive Disclosure

The 5 referenced files are clearly signaled one level deep ('See: references/X.md for...') and the body acts as a proper overview, but 6 of 11 bundle files (paradigm-elephant, paradigm-examples, paradigm-first-principles, paradigm-hamming, paradigm-technology-cycle, worked-examples) are not referenced from the body or any main reference and are therefore unnavigable — a notable navigation defect too large to call minor; not 2 because the signaled references are genuinely clear and well-structured.

3 / 5

Total

16

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is exemplary: third-person voice, concrete actions, an explicit 'Use when' clause with natural trigger phrases, and a distinct research-idea niche. It cleanly matches the top anchor on all four dimensions.

DimensionReasoningScore

Specificity

Names multiple concrete actions — 'five-dimension framework (Higher, Faster, Stronger, Cheaper, Broader)', 'idea-lifecycle and student-capability matching', 'paradigm-shift probing', 'fatal-flaws audit', 'reviewer-style verdict', and 'non-STEM ideas route to substitute frameworks' — giving comprehensive coverage of the skill's behavior rather than vague abstraction.

5 / 5

Completeness

Explicitly answers both 'what' (the evaluation framework and verdict output) and 'when' via a clear 'Use when the user has a draft research idea and asks...' clause with concrete trigger phrases, matching the anchor-5 pattern.

5 / 5

Trigger Term Quality

Surfaces the exact natural phrases a user would say — 'evaluate this idea', 'score this idea', 'assess feasibility', 'novelty check', 'is this a good research direction', 'whether it is worth pursuing', 'before committing to a paper scope' — with several synonyms covering the common request forms.

5 / 5

Distinctiveness Conflict Risk

Occupies a clear niche — evaluating preliminary research ideas for a reviewer-style verdict — with distinct research-idea triggers and explicit non-STEM routing, so it is unlikely to fire for unrelated skills.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
HKUSTDial/Supervisor-Skills
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.