CtrlK
BlogDocsLog inGet started
Tessl Logo

research-pipeline

Full end-to-end research pipeline: from a broad research direction through idea discovery, experiments, and review all the way to a polished paper PDF. Use when user says "全流程", "full pipeline", "从找idea到投稿", "end-to-end research", or wants the complete autonomous research lifecycle.

62

Quality

73%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skills-codex/research-pipeline/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

63%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A dense, highly actionable orchestration document: every stage has a concrete invocation command, gates, and outputs, with real validation and feedback loops. Its weaknesses are repetition (AUTO_PROCEED restated four times, a confusing Stage 5/6 numbering), a couple of dangling references, and heavy inline detail that belongs in reference files.

Suggestions

Consolidate the AUTO_PROCEED semantics into a single authoritative section and have Constants, Gate 1, and Key Rules reference it in one line each; delete the Stage 5/Stage 6 numbering explanation and just call it Stage 5.

Move the run-state resolution bash block, the run_state.py/resume/accept command details, and the watchdog/iteration-log protocol into a reference file (e.g. resumable-runs.md) and keep only the invocation summary in SKILL.md.

Replace the dangling "Write a final research status report (same as before)" with the actual required content, and either specify the /monitor-experiment [server] argument or drop the line.

DimensionReasoningScore

Conciseness

AUTO_PROCEED semantics are restated in four separate sections (Constants, Checkpoint execution rule, Gate 1, Key Rules), the provisional/accepted distinction appears three times, and the Stage 5/Stage 6 numbering explanation ("it is numbered Stage 5 here because this consolidated pipeline counts the writing handoff...") is meta-commentary that adds no operational value. Mostly operational and non-generic, but noticeably tighten-able — anchor 3 rather than 4, and not 2 since nothing pads with concepts Claude already knows.

3 / 5

Actionability

Concrete and largely executable: per-stage invocation commands ("/idea-discovery \"$ARGUMENTS\" — AUTO_PROCEED: $AUTO_PROCEED"), a literal bash block resolving RUN_STATE/ITER_LOG/WATCHDOG, exact run_state.py command lines, and ready-made checkpoint message templates. Minor gaps keep it below anchor 5: "Write a final research status report (same as before)" is a dangling reference, and "/monitor-experiment [server]" is an unexplained placeholder.

4 / 5

Workflow Clarity

Clear Stage 1→5 sequence with selection gates, sanity-check-first validation ("runs the smallest experiment first... auto-debugs failures (up to 3 attempts)"), a bounded review feedback loop ("repeat until score ≥ 6/10 or 4 rounds reached"), explicit fail-gracefully rules, and a resume-state phase table. Validation is present so no batch-operation cap applies, but the dual "Stage 5 / Stage 6" numbering and the scattered AUTO_PROCEED rules introduce minor incoherence — anchor 4 rather than 5.

4 / 5

Progressive Disclosure

No bundle files exist, and ~50 lines of run-state/watchdog machinery plus full report templates are inlined in SKILL.md where a reference file would serve better. The references it does make ([external-cadence.md](../shared-references/external-cadence.md), [resumable-runs.md](../shared-references/resumable-runs.md), the output protocols) are clearly signaled and one level deep, and section structure is good — anchor 3 rather than 2 (structure and signaling are present) and short of 4 (substantial content that should be separate remains inline with no offloading bundle).

3 / 5

Total

14

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states the full pipeline concretely, includes an explicit bilingual trigger clause, and answers both what and when. The only weaknesses are a few missing English trigger variants and minor overlap risk with the component skills it orchestrates.

DimensionReasoningScore

Specificity

"from a broad research direction through idea discovery, experiments, and review all the way to a polished paper PDF" lists several concrete pipeline actions with a tangible end artifact. It is not a 5 because coverage has minor gaps (literature survey, ablations, writing sub-steps are not mentioned), and not a 3 because it names more than 1-2 concrete actions.

4 / 5

Completeness

It explicitly answers both questions: the "what" is the concrete pipeline ("idea discovery, experiments, and review all the way to a polished paper PDF") and the "when" is an explicit "Use when user says ..." clause with quoted trigger phrases. Anchor 5 fits; anchor 4's caveat (a vaguer "when") does not apply.

5 / 5

Trigger Term Quality

Triggers "全流程", "full pipeline", "从找idea到投稿", "end-to-end research", and "complete autonomous research lifecycle" give good natural coverage including bilingual synonyms and a colloquial phrase. Not a 5 because common English variants such as "from idea to paper", "autonomous research", or "write and submit a paper" are missing.

4 / 5

Distinctiveness Conflict Risk

"Full end-to-end research pipeline" carves a clear umbrella niche distinct from its component workflows, but triggers like "end-to-end research" carry minor overlap risk with closely related sub-skills (idea-discovery, auto-review-loop, paper-writing). Distinct but not minimal-conflict, so anchor 4 rather than 5.

4 / 5

Total

17

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

relative_links

Relative link issues: 5 suspicious

Warning

Total

15

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.