CtrlK
BlogDocsLog inGet started
Tessl Logo

cross-modal-review

Quality gate via second model. Spawn a different AI model to review work before committing. Includes refusal routing: if one model refuses, switch silently to the next. Extended in v0.25.1 with structured review-mode gating (when to invoke vs not) and a Codex code-review handoff for the diff-review case.

51

Quality

57%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

High

Do not use without reviewing

Fix and improve this skill with Tessl

tessl review fix ./skills/cross-modal-review/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

56%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is well-structured with clear phases, concrete gating thresholds, and exact output templates, and it correctly defers detail to reference files. But the central 'spawn a different model' mechanic is never made executable, all referenced convention files are absent from the bundle, and version history plus a conformance-test stub inflate the token budget without adding operational value.

Suggestions

Remove the terminal '## Output Format' conformance-test stub (or merge it into the real 'Output format' section) and relocate inline version references (v0.25.1, v0.27.x) to a changelog or deprecated section.

Make the 'Spawn review model' phase executable — either inline the model-selection pairs and the actual spawn invocation, or ship the referenced conventions files (model-routing.md, cross-modal.yaml) inside the skill's references/ directory so the links resolve.

Add an explicit failure-path loop after the Grade step (e.g., 'If ISSUES FOUND: present findings, await user decision, optionally re-review after fixes') so the workflow's validation checkpoint has a recovery path.

DimensionReasoningScore

Conciseness

The body is mostly efficient — tight Contract bullets, a gating section with concrete thresholds, and exact output templates — but contains unnecessary padding: version references scattered inline ("v0.25.1 gating", "v0.25.1 extension", "v0.27.x"), a long blockquote disambiguating "gbrain eval cross-modal", and a duplicate terminal "## Output Format" stub whose only stated purpose is to satisfy a conformance test. Matches 'mostly efficient but includes some unnecessary explanation or could be tightened'; not 4 because the version-history material is time-sensitive content outside any old-patterns/deprecated section and the conformance stub is pure redundancy.

3 / 5

Actionability

Concrete elements exist — exact output framing templates with delimiters and field layout, specific invoke/don't-invoke thresholds ("5+ files or 100+ lines", "2+ iterations"), and a numbered refusal-routing chain — but the core mechanic is underspecified: "Spawn review model. Send the work + Contract to a different model" with no commands or mechanism, deferring model selection to conventions/model-routing.md which is not in the bundle. Matches 'some concrete guidance but incomplete... missing key details'; not 4 because the central execution step cannot be carried out from what is written here.

3 / 5

Workflow Clarity

The Phases section gives a clear five-step sequence (Capture → Load Contract → Spawn → Grade → Report) with a built-in checkpoint ("Grade... Pass / fail with specific citations") and the refusal-routing section adds an escalation fallback ("If ALL models in the chain refuse, escalate to the user"), with the user-sovereignty rule closing the loop. Matches 'clear sequence with most checkpoints present; minor validation gaps'; not 5 because there is no explicit feedback loop for what happens when the grade fails (e.g., rework and re-review) — only reporting to the user.

4 / 5

Progressive Disclosure

The body has real section structure and its references (../conventions/cross-modal.yaml, ../conventions/test-before-bulk.md, ../conventions/model-routing.md, skills/testing/SKILL.md) are one level deep and clearly signaled — but none of these files exist in the bundle (no references/, scripts/, or assets/ directories), so the links point outside the skill and cannot be navigated from it, and all substantive content lives inline in a ~175-line monolithic body capped by a redundant conformance-test section stub. Matches 'some structure but could be better organized; references present but not clearly signaling resolvable destinations'.

3 / 5

Total

13

/

20

Passed

Description

58%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description clearly conveys the skill's purpose and lists several concrete capabilities, but it reads partly like a changelog and entirely omits explicit 'when to use this' guidance and the natural trigger phrases users would say (those live only in the separate triggers list). Distinctiveness is good, with minor overlap risk against general code-review skills.

Suggestions

Add an explicit 'Use when...' clause to the description itself stating concrete invocation conditions (e.g., 'Use when the user asks for a second opinion, a double-check of significant code changes, or an adversarial review before commit').

Fold natural trigger phrases users would actually say — 'second opinion', 'double-check this', 'get another perspective' — into the description text rather than leaving them only in the triggers frontmatter list.

Move version-history sentences ('Extended in v0.25.1 with...') out of the description into the body or a changelog, and use the freed space to name the grading mechanism (review against the originating skill's Contract).

DimensionReasoningScore

Specificity

The description names several concrete actions — "Spawn a different AI model to review work before committing", "refusal routing: if one model refuses, switch silently to the next", "structured review-mode gating (when to invoke vs not)", and "a Codex code-review handoff for the diff-review case" — which matches the anchor for listing several specific actions with minor gaps (e.g., grading against the originating skill's Contract and output shape are not mentioned). Not 5 because coverage is not comprehensive, and the changelog-style clause "Extended in v0.25.1..." pads rather than describes a capability.

4 / 5

Completeness

The 'what' is clear (spawn a different model to review work pre-commit, with refusal routing and a Codex handoff), but the 'when' is only gestured at — "structured review-mode gating (when to invoke vs not)" describes that gating exists without stating any actual invocation conditions. Per the guideline, a missing 'Use when...' clause or equivalent explicit trigger guidance caps completeness at 3; not 4 because no concrete 'when' is given in the description.

3 / 5

Trigger Term Quality

The description text contains some relevant keywords ("review work before committing", "second model", "refuses") but omits the natural phrases a user would actually say — "second opinion", "double-check this", "get another perspective", "adversarial review" — which exist only in the separate `triggers` frontmatter list, not in the description itself. Matches the anchor 'some relevant keywords but missing common variations or synonyms'; not 4 because the natural synonyms users would voice are largely absent from the evaluated text.

3 / 5

Distinctiveness Conflict Risk

"Quality gate via second model", "refusal routing", and "Codex code-review handoff" carve out a fairly distinct niche (cross-model second opinion) with minimal overlap risk; the main residual overlap is with generic code-review skills via "review work before committing". Mostly distinct with minor overlap risk against closely related review/testing skills; not 5 because the phrase could still collide with a plain code-review skill.

4 / 5

Total

14

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 3 suspicious

Warning

Total

14

/

16

Passed

Repository
garrytan/gbrain
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.