CtrlK
BlogDocsLog inGet started
Tessl Logo

research-refine

Turn a vague research direction into a problem-anchored, elegant, frontier-aware, implementation-oriented method plan via iterative GPT-6-Astra review. Use when the user says "refine my approach", "帮我细化方案", "decompose this problem", "打磨idea", "refine research plan", "细化研究方案", or wants a concrete research method that stays simple, focused, and top-venue ready instead of a vague or overbuilt idea.

67

Quality

81%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

77%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An exceptionally actionable, well-sequenced workflow skill: exact prompts, templates, checkpoints, and feedback loops leave nothing ambiguous. Its weaknesses are length and duplication (principles, dimensions, and report templates restated multiple times) plus a fully monolithic structure that inlines large templates instead of splitting them into reference files.

Suggestions

Move the round-0 proposal template, the reviewer prompt, and the Phase 5 report templates into reference files (e.g. references/proposal-template.md, references/reviewer-prompt.md) and link them one level deep, cutting SKILL.md to the workflow overview plus phase instructions.

Delete the 'Key Rules' restatements that duplicate the four Overview principles and the phase-flow chart, and consolidate REVIEW_SUMMARY.md and REFINEMENT_REPORT.md templates, which share the score-evolution table and round-by-round record.

State the 7 reviewer scoring dimensions once (in the reviewer prompt) and reference them from Step 3.1 and the table headers instead of re-enumerating them in each location.

DimensionReasoningScore

Conciseness

The 721-line body is fully operational (no explanations of concepts Claude already knows), but contains real redundancy: the four principles appear in 'Overview' and again nearly verbatim in 'Key Rules', the phase-flow chart duplicates the Workflow prose, the 7 scoring dimensions are enumerated three-plus times, and the REVIEW_SUMMARY and REFINEMENT_REPORT templates substantially overlap. This fits 'mostly efficient but could be tightened' rather than the noticeably-verbose anchor, since nearly all padding is duplication of genuinely operational content.

3 / 5

Actionability

Guidance is fully executable: named constants (MAX_ROUNDS=5, SCORE_THRESHOLD=9, OUTPUT_DIR), a copy-paste-ready reviewer prompt with explicit dimension weights and verdict rules, exact markdown templates for every output file, checkpoint JSON snippets, and concrete resume/fresh-start decision rules.

5 / 5

Workflow Clarity

The Phase 0-5 sequence is explicit with per-phase checkpoints, a stop condition (score >= threshold, READY verdict, no drift), staleness rules for checkpoint recovery, and a genuine review-revise-re-evaluate feedback loop with anchor and simplicity checks before each revision.

5 / 5

Progressive Disclosure

No bundle files exist (references/, scripts/, assets/ are absent), yet the 721-line body inlines the 86-line proposal template, the ~70-line reviewer prompt, and three large report templates that clearly belong in separate reference files, matching the 'content that should be separate is inline' anchor. Section headers are well-organized and the three shared-references links are clearly signaled, but they point outside the skill directory (../shared-references/, ../../shared-references/) and cannot be verified as part of this bundle, keeping this below the good-structure anchor.

3 / 5

Total

16

/

20

Passed

Description

86%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description with excellent trigger coverage (bilingual synonyms) and an explicit what/when structure. Its main weakness is a buzzword-heavy adjective chain that pads the single core action instead of enumerating concrete capabilities.

Suggestions

Replace stacked adjectives ('problem-anchored, elegant, frontier-aware, implementation-oriented') with 2-3 distinct concrete actions, e.g. 'freeze a problem anchor, draft a focused method proposal, and iteratively revise it against reviewer scores'.

Trim or drop the generic trigger 'decompose this problem', which overlaps with general problem-decomposition skills and dilutes the niche.

Cut the trailing qualifier clause ('stays simple, focused, and top-venue ready instead of a vague or overbuilt idea') — it restate the 'what' without adding actionable or trigger information.

DimensionReasoningScore

Specificity

The description names one core concrete action ("Turn a vague research direction into a ... method plan via iterative GPT-6-Astra review") but the adjective chain ("problem-anchored, elegant, frontier-aware, implementation-oriented, top-venue ready") adds flavor rather than additional specific actions, matching the '1-2 concrete actions' anchor rather than the 'several specific actions' anchor.

3 / 5

Completeness

It explicitly answers both what the skill does (turns a vague direction into an implementation-oriented method plan via iterative review) and when to use it, with a "Use when the user says..." clause listing concrete trigger phrases.

5 / 5

Trigger Term Quality

Trigger coverage is comprehensive with natural user phrasings and their synonyms in both English and Chinese ("refine my approach", "refine research plan", "decompose this problem", "帮我细化方案", "打磨idea", "细化研究方案"), plus a descriptive catch-all condition.

5 / 5

Distinctiveness Conflict Risk

The research-method-refinement niche is clear and triggers are specific, but "decompose this problem" is a generic phrase and the skill family (research-refine-pipeline, idea-creator, experiment-plan) creates minor overlap risk with closely related skills, matching the 'mostly distinct; minor overlap risk' anchor.

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (722 lines); consider splitting into references/ and linking

Warning

relative_links

Relative link issues: 4 suspicious

Warning

Total

14

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.