CtrlK
BlogDocsLog inGet started
Tessl Logo

research-refine

Turn a vague research direction into a problem-anchored, elegant, frontier-aware, implementation-oriented method plan via iterative GPT-6-Astra review. Use when the user says "refine my approach", "帮我细化方案", "decompose this problem", "打磨idea", "refine research plan", "细化研究方案", or wants a concrete research method that stays simple, focused, and top-venue ready instead of a vague or overbuilt idea.

68

Quality

83%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable and the workflow is exceptionally well-sequenced with validation, checkpointing, and recovery logic. Its weaknesses are redundancy (principles repeated three times, overlapping Phase 5 report templates) and the absence of any progressive disclosure — large templates and the full reviewer rubric are inlined in SKILL.md rather than split into reference files.

Suggestions

Move the ~90-line proposal template (Step 1.6) and the reviewer rubric text (Phase 2/4 bundles) into references/ files (e.g. references/proposal-template.md, references/reviewer-prompt.md) and link to them, keeping SKILL.md as an overview.

State the four core principles once in the Overview and remove their verbatim repetition in Step 1.4 and Key Rules, keeping only any rule not already covered.

Consolidate the three overlapping Phase 5 report templates: REVIEW_SUMMARY.md, REFINEMENT_REPORT.md, and score-history.md duplicate the score-evolution and round-by-round tables; define each table once and reference it.

DimensionReasoningScore

Conciseness

The core principles ("One paper, one dominant contribution", "The smallest adequate mechanism wins", modern-leverage-as-prior) are restated three times across Overview, Step 1.4, and Key Rules, and the Phase 5 report templates (REVIEW_SUMMARY.md, REFINEMENT_REPORT.md, score-history.md) overlap substantially with each other. The templates are actionable rather than filler, but the repetition and template overlap mean the body could be meaningfully tightened — anchor 3, not 4 (which requires only minor trims).

3 / 5

Actionability

The body provides exact output file paths, a concrete JSON state schema with a field-definition table, explicit MCP invocation blocks (model, config, prompt shape), and full copy-paste markdown templates for every output artifact — an engineer could execute the workflow verbatim. This matches anchor 5 (fully executable, copy-paste ready, covers common cases).

5 / 5

Workflow Clarity

The phased workflow (0-5) is clearly sequenced with an explicit review-revise loop, a precise stop condition ("Overall score >= SCORE_THRESHOLD and verdict is READY and no unresolved drift"), and checkpoint-recovery logic including a resume-phase table and stale-state rules. Validation is embedded throughout (Anchor Check, Simplicity Check, drift warnings, score threshold) — anchor 5 with explicit validation and feedback loops.

5 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are all absent), yet the body inlines roughly 90 lines of proposal template plus the full reviewer rubric and re-evaluation bundle text — content that clearly belongs in references/ files. Section headers and the shared-references links are well signaled, so structure is present but significant content that should be separate is inline — anchor 3 rather than 4.

3 / 5

Total

16

/

20

Passed

Description

90%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with excellent trigger-term coverage (natural phrases in both English and Chinese) and an explicit what/when structure. The only weakness is mild buzzword padding in the adjective stack, which slightly dilutes the concreteness of the stated capability.

DimensionReasoningScore

Specificity

The description names the domain and one composite concrete action ("Turn a vague research direction into a ... method plan via iterative GPT-6-Astra review") but does not list several distinct capabilities, and the stacked adjectives ("problem-anchored, elegant, frontier-aware, implementation-oriented") lean toward buzzword padding rather than concrete actions. It matches anchor 3 (domain plus 1-2 concrete actions) rather than anchor 4 (lists several specific actions).

3 / 5

Completeness

It explicitly answers both questions: what ("Turn a vague research direction into a problem-anchored ... method plan via iterative GPT-6-Astra review") and when ("Use when the user says ... or wants a concrete research method ... instead of a vague or overbuilt idea"), with concrete trigger phrases. This matches anchor 5 exactly.

5 / 5

Trigger Term Quality

Trigger coverage is comprehensive with natural synonyms in two languages — "refine my approach", "decompose this problem", "refine research plan", "帮我细化方案", "细化研究方案", "打磨idea" — plus a situational description ("wants a concrete research method that stays simple, focused, and top-venue ready"). This matches anchor 5's comprehensive synonym coverage; nothing common is missing.

5 / 5

Distinctiveness Conflict Risk

It carves a clear niche (research method-plan refinement via an iterative external-review loop) with distinct bilingual triggers, and the GPT-6-Astra review mechanism plus the 'vague vs overbuilt idea' framing clearly separate it from adjacent research-planning or experiment-execution skills. Anchor 5 (clear niche, distinct triggers, minimal conflict risk) is the best fit.

5 / 5

Total

18

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (771 lines); consider splitting into references/ and linking

Warning

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

relative_links

Relative link issues: 3 suspicious

Warning

Total

13

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.