CtrlK
BlogDocsLog inGet started
Tessl Logo

figure-composer

Compose one publication-grade multi-panel figure. Entry from a one-line claim + data files, OR from an existing figure via `derive_outline_prompt` (you read the PNG). Runs a per-figure loop: outline (12-col grid, per-panel ask + label_budget) → render each panel with `panel_task` (loading `figure-style`), one at a time or parallelized → tile + stamp letters with `compose_figure` → adversarial composite self-review with two-tier feedback (Tier-1 outline_revisions / Tier-2 per-panel violations) → regen affected panels, ≤3 rounds. Helpers: panel_task / compose_figure / compose_crops / composite_review_task / derive_outline_prompt. For one standalone plot use `figure-style`; for whole-paper figure ordering use `paper-narrative`.

64

Quality

75%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./configs/microservice/bff-service/configs/agent-skills/claude-science/figure-composer/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong operational skill body: the workflow is clearly sequenced with explicit validation checkpoints, bounded feedback loops, and concrete code for the common path. The main weaknesses are repetition of the figure-style-loading and no-API points, and helper contracts (review_schema, rules_path, min_floor) that are referenced but never specified, forcing reliance on the not-included kernel.py.

Suggestions

State the `figure-style` loading requirement once (Step 0) and reference it elsewhere with a single phrase instead of re-explaining it in §0 and §2; the same applies to the "no API, your own judgment" clarification.

Specify the `review_schema()` shape (the `editor_verdict`, `outline_revisions`, and `violations` fields the loop reads) the way the outline schema is shown in §1, since the §4 loop depends on it.

Include the kernel.py (or at least its helper signatures and the semantics of `rules_path`, `min_floor`, and `fig_label`) in the bundle so the SKILL.md's central references are verifiable.

DimensionReasoningScore

Conciseness

The body is dense and operational with no filler, but key points repeat: the requirement to load `figure-style` is stated in Step 0, §0, and again in §2 ("each runs its `panel_task` prompt, loads `figure-style` itself"), and "this is your own judgment, not an API call" / "no host runtime and no LLM API" each appear twice. This matches anchor 4 (efficient with minor over-explanation that could be trimmed); not 3 because there is no padded explanation of concepts Claude already knows.

4 / 5

Actionability

Mostly executable: the exec line for kernel.py, `pip install pillow matplotlib`, a concrete JSON outline example, and runnable snippets (`panel_task(...)` dict comprehension, the PIL crop loop in §3.5). Gaps: the §4 loop block is structured pseudocode rather than executable code, and helper arguments like `rules_path`, `min_floor`, `fig_label`, and the `review_schema()` fields are named but not specified, so the reader must consult kernel.py. Fits anchor 4 (mostly executable with minor gaps); not 5 because the core review loop is not copy-paste ready.

4 / 5

Workflow Clarity

Exemplary sequence (Step 0 → entry points → §1 outline → §2 render → §3 compose → §3.5 vision QA → §4 adversarial loop) with explicit validation checkpoints and feedback loops: a pre-review crop inspection pass, a bounded loop ("max 3 rounds, floor 5→4→3"), explicit break conditions ("editor_verdict in {accept, minor_revision} and 0 BLOCKER and ≤2 MAJOR"), regression tracking via `prev_path`, regenerating only affected panels, a convergence signal, and an anti-patterns section. Matches anchor 5.

5 / 5

Progressive Disclosure

Sections are well-organized and references are one level deep and clearly signaled (`figure-composer/kernel.py` for code, `figure-style` for design rules, `paper-narrative` for paper-level ordering), with the heavy lifting correctly delegated to kernel.py rather than inlined. However, the bundle contains no references/, scripts/, or assets/ directories, so the central `kernel.py` dependency cannot be verified as a real file — a minor organization gap that keeps this at anchor 4 rather than 5.

4 / 5

Total

17

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A highly specific, well-disambiguated description that clearly states what the skill does, but it buries the trigger conditions: there is no explicit "Use when..." clause for this skill itself (only routing guidance for its siblings), and the second-person "(you read the PNG)" costs it specificity. The helper/jargon density ("label_budget", "Tier-1/Tier-2") also crowds out natural user phrasing.

Suggestions

Add an explicit trigger clause, e.g. "Use when the user wants a single publication-quality multi-panel (composite) scientific figure from data files or an existing figure."

Replace second-person phrasing like "(you read the PNG)" with third person ("the agent reads the PNG") per description voice conventions.

Trade some internal jargon ("label_budget", "Tier-1 outline_revisions / Tier-2 per-panel violations") for natural synonyms users would say: "scientific figure", "composite figure", "panel labels", ".png".

DimensionReasoningScore

Specificity

Lists multiple concrete actions ("Compose one publication-grade multi-panel figure", "render each panel with `panel_task`", "tile + stamp letters with `compose_figure`", "adversarial composite self-review", "regen affected panels, ≤3 rounds") matching the comprehensive anchor 5, but "(you read the PNG)" is second-person voice, which the guidelines penalize by reducing specificity by 1.

4 / 5

Completeness

The "what" is clear and detailed, but there is no explicit "Use when..." clause for this skill; the "when" is only weakly implied through routing sentences about sibling skills ("For one standalone plot use `figure-style`; for whole-paper figure ordering use `paper-narrative`"). Per the judging guidelines, a missing 'Use when...' clause or equivalent caps completeness at 3, matching the anchor "clear 'what' but 'when' missing or only weakly implied".

3 / 5

Trigger Term Quality

Good natural-term coverage ("multi-panel figure", "publication-grade", "plot", "figure") but leans on internal jargon ("label_budget", "Tier-1 outline_revisions / Tier-2 per-panel violations", "12-col grid") and misses common synonyms/extensions like "scientific figure", "composite figure", or ".png". Not 3 because several genuinely natural phrases are present; not 5 because synonyms and file extensions are absent.

4 / 5

Distinctiveness Conflict Risk

Clear niche (composing one multi-panel publication figure) with explicit disambiguation from the two adjacent skills ("For one standalone plot use `figure-style`; for whole-paper figure ordering use `paper-narrative`"), which minimizes wrong-skill triggering. Matches the anchor for a clear niche with distinct triggers.

5 / 5

Total

16

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
UnicomAI/wanwu
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.