Evaluates a preliminary research idea against a five-dimension framework (Higher, Faster, Stronger, Cheaper, Broader) plus idea-lifecycle and student-capability matching, paradigm-shift probing, and a fatal-flaws audit. Returns a reviewer-style verdict; non-STEM ideas route to substitute frameworks. Use when the user has a draft research idea and asks whether it is worth pursuing, asks to 'evaluate this idea', 'score this idea', 'assess feasibility', 'novelty check', 'is this a good research direction', or before committing to a paper scope.
75
92%
Does it follow best practices?
Run evals on this skill
Adds up to 20 points to the overall score
View guide
Low
Low-risk findings worth noting
This skill evaluates a preliminary research idea from the combined perspective of a top-venue reviewer and an experienced advisor. It scores the idea against five improvement dimensions from the idea-generation guide (Higher, Faster, Stronger, Cheaper, Broader), matches the idea's lifecycle against the user's actual capability and available hours per week, probes whether the idea has paradigm-shift potential, flags fatal flaws, and returns one of three verdicts: Strong Accept, Accept with Revisions, or Reject and Pivot.
The goal is to kill weak ideas before the student invests months, and to shape promising-but-underdeveloped ideas into stronger forms before writing begins.
intro-drafter, tech-paper-template, or
benchmark-paper-template (separate plugin) instead.pre-submission-reviewer.benchmark-paper-template (separate plugin) in targeted mode.Read the user's idea description. In one paragraph, state whether the idea reads as Novel Problem, Novel Method, or New Setting. Is the story compelling in one sentence? If you cannot write that sentence, the idea itself is probably not yet clear enough for evaluation; ask the user to restate.
While restating, classify the research paradigm by method, not by department name: experiments, benchmarks, ablations, or model architectures mean STEM (continue with the five dimensions and fatal flaws below); text analysis, archives, or conceptual argument mean humanities; surveys, interviews, statistics, or fieldwork mean empirical social science; regressions, IV, DID, RDD, or panel data mean finance or economics; statutes, cases, or doctrine mean law.
See: references/domain-evaluation-frameworks.md for the substitute five-dimension frameworks and fatal-flaw substitutions used by the non-STEM paradigms. When the paradigm is unclear, ask one question: "is the core method experiments, surveys, text analysis, or theoretical derivation?"
See: references/fatal-flaws.md for the ten canonical fatal flaws, each with a detection rule and a defense strategy.
Run the fatal-flaws audit before the scoring steps rather than after them. Identify at most two fatal flaws. For each, state the flaw, cite the detection rule, and recommend a concrete defense.
Ground the novelty flaw (F1) in real retrieval whenever the environment has a literature-search capability (a scholarly search tool, web search over scholarly indexes, or shell access to public APIs). Extract two or three keyword groups from the idea (core method plus domain; mechanism plus task; technique plus benchmark), search, and name the three to five closest published works with title, authors, and year. For each, state which axis actually differs: the object acted on, the mechanism, the input granularity, or the problem setting. A similar title alone never establishes duplication; duplication requires failing to find even one differing axis. Retrieval results support metadata-level judgments only (who did what, where); never quote numbers or method details from search snippets. And "not found" does not prove novelty: report it as "no directly overlapping work retrieved under these keywords". With no retrieval capability, label the novelty judgment "unverified; literature check required".
A data-refuted core mechanism is an automatic CRITICAL. If the user's own reported data or attachments already show the core mechanism matched or beaten by a baseline or simple control, lock the verdict to Reject and Pivot; write no defense and invent no optimistic threshold. See: references/fatal-flaws.md, the data-refuted section, for the exact boundary between "refuted by data" and "merely untested".
Short-circuit rule. If any fatal flaw is tagged CRITICAL in the severity taxonomy (single-handedly causes rejection, unfixable within the lifecycle), stop here and emit the verdict directly:
If no CRITICAL flaw is found, continue to Step 3.
See: references/lifecycle-capability-matching.md for the six-category lifecycle matrix, capability self-assessment rubric, and mismatch recovery strategies.
Map the idea onto one of six categories (Application, Foundational Theory, Cross-Disciplinary, Frontier Exploration, Data-Intensive, Innovative Technique). Match against the user's declared capability (effective hours per week, skill depth, theoretical versus applied strength). Output a mismatch flag if lifecycle is shorter than the user's realistic execution window.
See: references/five-dimensions.md for each dimension's entry strategies, scoring rubric, and worked examples.
Score the idea on each of:
Score each 1-10 with explicit evidence from the user's stated contribution. Identify the two or three dimensions where the idea has the highest ceiling and recommend emphasising those in the paper.
Scoring discipline: start every dimension at 5 and justify movement. Two kinds of grounds move a score up, and both count: measured results the user reported (quote them), or a mechanism argument that holds up (label the score "mechanism-based, not yet confirmed by data"). A solid, untested mechanism can reach 8 or 9 with that label plus a named validation experiment; do not systematically cap untested ideas. A dimension with neither data nor mechanism stays at 5 with "no grounds given". Watch attribution: when an impressive gain plausibly comes from a peripheral factor (routing, post-processing, a stronger base model, favorable samples), cap that dimension until an ablation isolates the core mechanism. The scoring reference's final two sections cover both rules in detail.
For non-STEM paradigms, score the substitute dimensions from references/domain-evaluation-frameworks.md instead, under the same discipline.
See: references/paradigm-shift-probe.md for the four probing principles (First Principles, Elephant in the Room, Technology Cycle, Hamming's Rule) and the cross-reference to handbook section 2.3 when deeper disruptive-innovation exploration is needed.
Test the idea against four questions:
Two or more yes answers means the idea has disruptive potential. Note that, and recommend reading handbook 2.3 to deepen the thinking on disruptive-innovation dimensions.
Against the user's stated resources (hardware, data access, team size, engineering skills, timeline), assess:
If any risk is high, flag it explicitly with a suggested mitigation.
Before emitting the verdict, run the checks in the Integrity gate section below.
Issue one of three verdicts:
When the high scores are mechanism-based rather than data-based, qualify the verdict as "worth pursuing, pending the validation experiment", and name that experiment in the top-three actions.
Emit the evaluation in the Output format below.
Each bullet is tagged with an enforceability class. [inspection] means the LLM can verify the bullet from the produced output alone. [attestation] means the LLM states it has done the check, but the user remains responsible for verification. [user-attest] means the bullet is a user-side rule the skill cannot confirm.
Before returning the verdict:
If any [inspection] check fails, downgrade the verdict and mark the corresponding output section as "needs user attention". For [attestation] bullets, the skill states the check was run and the user confirms the result.
Run the gate silently. Do not print a per-gate pass or fail report; a failure surfaces as a concrete finding inside the affected output section, and the delivered evaluation stays free of internal checking rituals.
| # | Flaw | Severity | Defense |
|---|---|---|---|
| 1 | ... | CRITICAL or MAJOR | ... |
If any CRITICAL flaw is present, skip sections 3-6 and go to section 7 with verdict Reject and Pivot.
| Aspect | User's input | Assessment |
|---|---|---|
| Idea category | ... | ... |
| Lifecycle | ... months | ... |
| Weekly effective hours | ... | ... |
| Fit | ... | Green or Yellow or Red |
| Dimension | Score 1-10 | Evidence | Lift suggestion |
|---|---|---|---|
| Higher | ... | ... | ... |
| Faster | ... | ... | ... |
| Stronger | ... | ... | ... |
| Cheaper | ... | ... | ... |
| Broader | ... | ... | ... |
| Probe | Yes or No | Rationale |
|---|---|---|
| First Principles | ... | ... |
| Elephant in the Room | ... | ... |
| Technology Cycle | ... | ... |
| Hamming's Rule | ... | ... |
Disruptive potential: <none, possible, strong>.
| Risk | Level | Mitigation |
|---|---|---|
| Compute | ... | ... |
| Data | ... | ... |
| Engineering | ... | ... |
| Timeline | ... | ... |
(mechanism-based high scores: append "worth pursuing, pending the validation experiment")
Top three actions to take first:
aff5de9
If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.