CtrlK
BlogDocsLog inGet started
Tessl Logo

auto-paper-improvement-loop

Autonomously improve a generated paper via Claude review through claude-review MCP → implement fixes → recompile, for 2 rounds. Use when user says "改论文", "improve paper", "论文润色循环", "auto improve", or wants to iteratively polish a generated paper.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skills-codex-claude-review/auto-paper-improvement-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

52%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers highly actionable, well-validated workflow guidance with concrete commands, schemas, and prompt templates, but it is a monolithic document with no progressive disclosure. Pervasive duplication (the review prompt, reviewer-independence rules, and whitelist rules each appear multiple times) and a ~100-line opt-in spec inlined in the main file inflate token cost, and the duplicate 'Step 9' headings and tangled step numbering slightly impair navigation.

Suggestions

Split the Edit Whitelist specification (schema, glob semantics, detectors, behaviors — ~100 lines of opt-in detail) into a references/edit-whitelist.md file and keep only a short summary plus the invocation flags in SKILL.md, directly addressing progressive_disclosure.

Deduplicate the reviewer-independence rules and the review prompt template: state the rules once in the Reviewer Independence Protocol section, reference them from Constants/Steps/Key Rules, and define the prompt once with only the round-specific delta (e.g. the 'fresh, zero-context review' preamble) inlined for Round 2 — this is the biggest conciseness win.

Fix the two sections both numbered 'Step 9' (renumber the summary step, e.g. to Step 10) and straighten the step numbering (2b/4.5/5/5.5/5b) into a consistent sequence so the workflow is easier to follow.

DimensionReasoningScore

Conciseness

The body is noticeably verbose through large-scale duplication: the ~45-line reviewer prompt template appears nearly verbatim in both Step 2 and Step 5; the reviewer-independence rules are stated three times (REVIEWER_BIAS_GUARD constant, the Reviewer Independence Protocol section, and Key Rules); and the edit-whitelist rejection-logging rules are repeated in Constants, Steps 3/6, and Key Rules. It is not 1 because it never explains basic concepts Claude already knows, and not 3 because the redundancy is pervasive rather than occasional tightening opportunities.

2 / 5

Actionability

Mostly executable guidance: concrete bash commands (latexmk pipeline, pdfinfo page check, duplicate-label grep), full MCP prompt templates, YAML/JSON schemas, a regex detector table, and a state-file spec. Minor gaps keep it below 5 — the Step 4.5 Python heredoc is a skeleton whose comparison logic is only a comment, the format-check greps over the log are partially illustrative, and [VENUE]/[list figure files] placeholders must be filled. It is well above 3, where guidance would be pseudocode-level.

4 / 5

Workflow Clarity

Steps 0–9 are clearly sequenced with explicit validation checkpoints (verify 0 undefined references after recompile, the Step 4.5 restatement regression test, the Step 8 format check with hard stop criteria) plus state persistence for crash recovery. It falls short of 5 because two different sections are both numbered 'Step 9' (Document Results and Summary) and the step numbering is chaotic (2b, 4.5, 5, 5.5, 5b), which muddies an otherwise strong sequence. Validation is present throughout, so the missing-validation cap does not apply.

4 / 5

Progressive Disclosure

There are no bundle files at all — references/, scripts/, and assets/ are absent — and everything lives in this single ~620-line SKILL.md. Content that clearly belongs in separate files is inlined: the ~100-line opt-in Edit Whitelist specification, the duplicated reviewer prompt templates, and the detailed format-check rules. The only file reference ('../shared-references/review-tracing.md') points outside the skill directory and cannot be verified from the bundle. This fits the level-2 anchor better than level 3, since there is essentially no in-bundle reference structure to signal.

2 / 5

Total

12

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description that clearly and explicitly states both the multi-action pipeline and a rich set of natural trigger phrases, including bilingual synonyms. Its only weaknesses are minor: a few missing natural variations of the trigger terms and slight overlap risk with adjacent paper-writing skills on generic phrases like 'improve paper'.

Suggestions

Add one or two more natural trigger variations (e.g. 'polish paper', 'revise paper', 'paper review loop') to broaden keyword coverage for the trigger_term_quality dimension.

Sharpen distinctiveness by hinting at the scope boundary in the description itself (e.g. 'iteratively polish an already-compiled paper' vs. writing a paper from scratch) so generic 'improve paper' requests are less likely to collide with paper-writing skills.

DimensionReasoningScore

Specificity

The description lists multiple specific concrete actions covering the full pipeline — 'improve a generated paper via Claude review through claude-review MCP → implement fixes → recompile, for 2 rounds' — naming the tool, the loop stages, and the round count. This matches the comprehensive-coverage anchor; a score of 4 would require minor coverage gaps, which are not present.

5 / 5

Completeness

Explicitly answers both what ('improve a generated paper via Claude review... implement fixes... recompile, for 2 rounds') and when ('Use when user says "改论文", "improve paper"...') with concrete trigger phrases. This is a direct match for the 5 anchor; the 4 anchor would require the 'when' to be less explicit or specific.

5 / 5

Trigger Term Quality

Includes natural phrases users would say — '改论文', 'improve paper', '论文润色循环', 'auto improve', 'iteratively polish a generated paper' — with bilingual synonyms, but misses common variations such as 'revise paper', 'polish paper', or 'paper review'. Good coverage with a few natural terms missing fits the 4 anchor; it is not 5 because coverage is not comprehensive, and not 3 because several natural terms and synonyms are present.

4 / 5

Distinctiveness Conflict Risk

The iterative review→fix→recompile paper-polishing loop is a mostly distinct niche with specific triggers, but generic phrases like 'improve paper' carry minor overlap risk with general paper-writing or editing skills. This fits 'mostly distinct; minor overlap risk'; it is not 5 because some triggers are not uniquely bound to this skill, and not 3 because the loop framing and named triggers are fairly specific.

4 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (626 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.