CtrlK
BlogDocsLog inGet started
Tessl Logo

boltz-structure-and-binding

Predict structures and binding for one defined complex with Boltz. Use when folding a protein, RNA, DNA, or ligand complex, docking one ligand, predicting an interface, or scoring binding. Not for screening libraries or design.

73

Quality

90%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Critical

Do not install without reviewing

SKILL.md
Quality
Evals
Security

Quality

Content

81%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is highly actionable with a clearly sequenced workflow, cost confirmation before submit, error-recovery guidance, and well-signaled one-level-deep references that verifiably exist in the bundle. Its main weakness is redundancy: the Codex background/heartbeat/session handling is restated in the Workflow, the Command Pattern comments, and the Always-Do-This bullets, which both costs tokens and blurs the file's organization.

Suggestions

State the Codex background/heartbeat procedure once (e.g., in a short 'Runtime-specific download handling' section or a reference file) and reference it from Workflow step 6 and the Command Pattern comments instead of repeating it in three places.

Move the runtime-specific detail (Codex `session_id`/heartbeat polling, Claude Code permission-gating rules) into a small reference file, keeping only the one-sentence rule that applies to every run in the main body.

Group the 'Always Do This' bullets by theme (payload hygiene, paths, runtime/background handling, polling/cost) so related rules read as one block rather than an interleaved grab-bag.

DimensionReasoningScore

Conciseness

The body is dense and assumes Claude's competence (no explaining what folding or YAML is), but the Codex background/heartbeat handling is repeated nearly verbatim in three places — Workflow step 6 ("In Codex, run `download-results` as a foreground shell command with `yield_time_ms: 1000`; if Codex returns a `session_id`..."), the Command Pattern comments ("Codex: foreground shell command with yield_time_ms=1000; keep the returned session_id if one is provided"), and two long Always-Do-This bullets. This matches 'Mostly efficient but includes some unnecessary explanation or could be tightened'; the redundancy is more than the 'minor instances' of level 4.

3 / 5

Actionability

The Command Pattern block gives complete, executable `boltz-api` invocations with concrete flags (`--model boltz-2.1`, `--input @yaml:///absolute/path/payload.yaml`, `--raw-output --transform id`), and the payload shapes are shown as real JSON/YAML snippets with exact field names (`chain_ids`, `value`, `binder_chain_id`). Only placeholders the agent must fill (`<job-id-from-start>`, `<run-name>`) remain, and the file explicitly says to replace them — matching 'Fully executable; copy-paste ready code or commands'.

5 / 5

Workflow Clarity

The six-step Workflow is explicitly sequenced (normalize entities → binding block → author payload → `estimate-cost` → confirm → `start` → `download-results`) with a cost-confirmation checkpoint before the paid submit ("run `estimate-cost`, show the USD cost, wait for explicit confirmation") and error-recovery feedback: the 'SAB 400 validation quirk' section says "If the server rejects a payload ... inspect `entities`, `binding`, and `constraints`; read references/results.md for details", and restart guidance covers re-running `download-results` with the same `--name`. This matches 'Clear sequence with explicit validation steps; feedback loops for error recovery'.

5 / 5

Progressive Disclosure

Structure is good: the body is a workflow overview, and both referenced bundle files exist and hold what is claimed — references/api.md carries the payload field reference (verified: entity types, binding variants, bonds, constraints, model_options) and references/results.md carries output layout, nested metrics, and the validation quirk. All references are one level deep and clearly signaled with markdown links. It falls short of level 5 because the runtime-specific operational detail (~25 lines of Codex session_id/heartbeat and Claude Code permission-gating guidance spread across three sections) is inlined in SKILL.md rather than split into a reference with a one-line pointer — 'most content appropriately placed; minor organization gaps'.

4 / 5

Total

17

/

20

Passed

Description

100%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is a model frontmatter description: it states concrete capabilities in third person, provides explicit and natural 'Use when' triggers across all supported molecule types, and adds a negative boundary clause that cleanly separates it from screening/design skills. No dimension shows a defect against the top anchors.

DimensionReasoningScore

Specificity

"Predict structures and binding for one defined complex with Boltz" names concrete actions, and the trigger list enumerates specific operations — "folding a protein, RNA, DNA, or ligand complex, docking one ligand, predicting an interface, or scoring binding" — giving comprehensive coverage of the skill's domain. This matches the anchor 'Lists multiple specific concrete actions; comprehensive coverage'; there are no vague filler phrases, so it is not the level 4 example with coverage gaps.

5 / 5

Completeness

The 'what' is explicit ("Predict structures and binding for one defined complex with Boltz") and the 'when' is explicit with concrete trigger phrases ("Use when folding a protein, RNA, DNA, or ligand complex, docking one ligand, predicting an interface, or scoring binding"), matching the anchor 'Clearly and explicitly answers both what AND when with concrete trigger phrases'. The voice is third person ('Predict structures'), so no voice penalty applies.

5 / 5

Trigger Term Quality

The triggers are the natural phrases a user would say when they need this skill: "folding a protein", "docking one ligand", "predicting an interface", "scoring binding". Synonyms and variations are well covered across molecule types (protein, RNA, DNA, ligand), matching the anchor 'Comprehensive coverage of natural terms including synonyms'. It is above level 4 because no common way of asking for structure prediction or binding analysis is obviously missing.

5 / 5

Distinctiveness Conflict Risk

"for one defined complex" carves out a clear niche versus library-scale workflows, and the explicit negative boundary "Not for screening libraries or design" prevents mis-triggering against sibling screening/design skills. This matches 'Clear niche with distinct triggers; minimal conflict risk' and is stronger than the level 4 'minor overlap risk' example.

5 / 5

Total

20

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
openai/plugins
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.