CtrlK
BlogDocsLog inGet started
Tessl Logo

auto-paper-improvement-loop

Autonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds. Use when user says "改论文", "improve paper", "论文润色循环", "auto improve", or wants to iteratively polish a generated paper.

63

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./skills/auto-paper-improvement-loop/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

70%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Operationally excellent content — executable commands, explicit validation gates, and regression tests — but it is a monolithic 695-line file that inlines substantial spec material that belongs in reference files, and it carries notable redundancy (duplicated reviewer prompt, rationale prose). Splitting opt-in sections (edit whitelist, style-ref, kill-argument, reviewer prompt template) into references/ would fix both the progressive-disclosure and conciseness weaknesses at once.

Suggestions

Move the ~100-line Edit Whitelist spec, the Style Reference contract, and the reviewer prompt template into separate files under references/ (e.g. references/edit-whitelist.md, references/reviewer-prompt.md) and keep one-line pointers plus defaults in SKILL.md — the prompts are duplicated verbatim between Step 2 and Step 5 and only the bias-guard delta needs to be inline.

Tighten or relocate the "Rationale" and "Empirical motivation" prose paragraphs (e.g. under Step 4.5, Step 5.5, Step 8) into a single 'why these rules exist' note or drop them; they justify design choices but cost context on every invocation.

Fix the duplicate step numbering (two 'Step 9' headings) and make the Step 4.5 restatement check fully executable — the inline Python currently stops at the normalize() function with the comparison left as a comment.

DimensionReasoningScore

Conciseness

Mostly skill-specific knowledge rather than concepts Claude already knows, but the ~50-line reviewer prompt is duplicated nearly verbatim in Steps 2 and 5, "Rationale"/"Empirical motivation" prose is padded, and the ~100-line edit-whitelist spec plus full prompt templates are inlined, so it could be noticeably tightened (not a 2 — the bulk is non-redundant operational detail).

3 / 5

Actionability

Concrete, executable guidance dominates: exact bash commands (latexmk, pdfinfo, dup-label grep), full MCP call configs with model and reasoning-effort, a YAML whitelist schema, regex detectors, and invocation examples. Minor gaps: the Step 4.5 Python snippet ends in comments ("Compare normalized theorem blocks..." — the actual comparison logic is pseudocode), and [VENUE]/[list figure files] placeholders must be filled by hand.

4 / 5

Workflow Clarity

Steps 0–9 are clearly sequenced with explicit validation checkpoints and feedback loops: recompile with "Verify: 0 undefined references, 0 undefined citations", the Step 4.5 restatement regression test rerun after every recompile, format-check stop criteria with severity thresholds, and auto-fix tables. Only slip is cosmetic: two steps are both labeled "Step 9" (Document Results / Summary), which is a labeling defect, not a missing validation.

5 / 5

Progressive Disclosure

Section headers and in-body navigation are good, but the skill is a single 41KB monolithic file with no references/ bundle at all — the edit-whitelist spec, style-ref contract, restatement-check details, and both full reviewer prompts clearly belong in separate reference files ("content that should be separate is inline"). Cross-references (shared-references/*, tools/extract_paper_style.py, skills/kill-argument/SKILL.md) point outside the bundle and are at least clearly signaled, keeping this above the unstructured-inline anchor of 2.

3 / 5

Total

15

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: concrete what, explicit bilingual when-triggers, and a clearly scoped niche. Only minor gaps — a few missing English synonyms and slight overlap risk with closely related research/paper-loop skills.

DimensionReasoningScore

Specificity

Quotes "review → implement fixes → recompile, for 2 rounds" and "GPT-6-Astra xhigh review" — several concrete, specific actions in a named domain (generated papers). Minor gaps: the logging, format check, and kill-argument phases are not mentioned, so coverage is not comprehensive (not a 5).

4 / 5

Completeness

Explicitly answers both: what — "Autonomously improve a generated paper via GPT-6-Astra xhigh review → implement fixes → recompile, for 2 rounds" — and when — "Use when user says '改论文', 'improve paper', '论文润色循环', 'auto improve', or wants to iteratively polish a generated paper". Both are concrete with explicit trigger phrases, matching the top anchor.

5 / 5

Trigger Term Quality

Quotes natural triggers: "improve paper", "auto improve", "改论文", "论文润色循环", "iteratively polish a generated paper" — good bilingual coverage. A few natural English variations (e.g., "polish paper", "paper revision/refinement") are missing, so below the comprehensive-synonym anchor of 5.

4 / 5

Distinctiveness Conflict Risk

The niche (post-compilation paper-polish loop with external LLM review) is distinct and the triggers are loop-specific. Minor overlap risk: "improve paper" could collide with a sibling skill like the referenced `/auto-review-loop` or generic paper-editing skills.

4 / 5

Total

17

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (696 lines); consider splitting into references/ and linking

Warning

allowed_tools_field

'allowed-tools' contains unusual tool name(s)

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 1 suspicious

Warning

Total

12

/

16

Passed

Repository
wanshuiyin/Auto-claude-code-research-in-sleep
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.