CtrlK
BlogDocsLog inGet started
Tessl Logo

auto-paper-improvement-loop

Autonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says "改论文", "improve paper", "论文润色循环", "auto improve", or wants to iteratively polish a generated paper.

64

Quality

79%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/skills-codex/auto-paper-improvement-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an unusually rigorous, well-validated loop with concrete commands, prompt templates, stop criteria, and crash recovery — workflow clarity is exemplary. Its weaknesses are token efficiency (large verbatim duplication) and the total absence of progressive disclosure: 619 monolithic lines where reference files should carry the whitelist spec and prompt templates.

Suggestions

Extract the edit-whitelist specification (schema, resolution rules, detectors) into references/edit-whitelist.md and the reviewer prompt template into references/reviewer-prompt.md, leaving one-line pointers in SKILL.md — this also removes the Step 2/Step 5 verbatim duplication.

Replace the duplicated Steps 3/6 whitelist-gate prose with a single referenced procedure and fold the 'Empirical motivation' anecdotes into a short rationale line each.

Finish the Step 4.5 Python snippet so it actually compares normalized theorem blocks (it currently ends in comments), and fix the duplicate 'Step 9' section numbering.

DimensionReasoningScore

Conciseness

The ~50-line reviewer prompt is duplicated verbatim in Steps 2 and 5, the edit-whitelist gate prose is restated in Steps 3 and 6, and long 'Empirical motivation' paragraphs pad the 619-line body — content is operational rather than educational, but it could be materially tightened.

3 / 5

Actionability

Concrete latexmk/grep/pdfinfo commands, full spawn_agent prompt templates, and exact log/state-file schemas make the guidance mostly executable, but the Step 4.5 Python snippet is a non-executable skeleton ending in comments, and Step 2's placeholders ([VENUE], [list figure files]) require assembly.

4 / 5

Workflow Clarity

Steps 0–9 are clearly sequenced with explicit validation checkpoints (0 undefined references, the restatement regression test after every recompile, location-aware format-check stop criteria, JSON state recovery), severity-ranked fix priorities, and recompile-verify feedback loops — the only blemish is cosmetic: two sections are both numbered 'Step 9'.

5 / 5

Progressive Disclosure

The skill has no bundle files at all: the ~100-line edit-whitelist guide, the two prompt templates, and the fix-pattern tables are inlined in SKILL.md where they belong in one-level-deep reference files; internal section structure and external skill references are good, which keeps this above the minimal-structure anchor.

3 / 5

Total

15

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it names the exact pipeline, the reviewer model, the round count, and gives explicit bilingual trigger phrases with a clear 'Use when' clause. Its only weaknesses are a slightly thin synonym set and the somewhat generic 'auto improve' trigger that risks collision with other improvement skills.

Suggestions

Add missing natural trigger synonyms such as 'polish paper', 'revise paper', or 'paper review loop' to broaden match coverage toward the top anchor.

Qualify the generic 'auto improve' trigger (e.g. 'auto improve paper') to reduce conflict risk with code- or general-improvement skills.

DimensionReasoningScore

Specificity

The description states the concrete action pipeline 'GPT-6-Astra xhigh review → implement fixes → recompile', the iteration count ('for 2 rounds'), and the target artifact ('a generated paper'), giving multiple specific concrete actions with comprehensive coverage of the loop.

5 / 5

Completeness

It explicitly answers 'what' (review → implement fixes → recompile for 2 rounds on a generated paper) and 'when' via a literal 'Use when user says...' clause with concrete trigger phrases, matching the top anchor exactly.

5 / 5

Trigger Term Quality

Trigger phrases are natural and varied ('improve paper', '改论文', '论文润色循环', 'auto improve', 'iteratively polish a generated paper'), but common variations such as 'polish paper', 'revise paper', or 'paper review loop' are missing, so coverage is good rather than comprehensive.

4 / 5

Distinctiveness Conflict Risk

The paper-polishing niche is clear and mostly distinct, but the trigger 'auto improve' is generic and could fire for code-improvement requests, leaving minor overlap risk with sibling improvement skills.

4 / 5

Total

18

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (619 lines); consider splitting into references/ and linking

Warning

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

13

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.