CtrlK
BlogDocsLog inGet started
Tessl Logo

llava

Vision-language chat: VQA, captioning, image dialogue.

47

Quality

51%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./optional-skills/mlops/llava/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers a strong, runnable quick start and good model-selection tables, but roughly a third of the snippets depend on undefined helper functions, quantization content is duplicated, and the existing references/training.md bundle file is never referenced — leaving structure and navigation as the clearest weak points. Tightening duplication and either linking or deleting the orphaned reference would raise conciseness and progressive disclosure together.

Suggestions

Replace the undefined 'ask(model, image, question)' and 'generate(conv, model, image)' helpers with the real generation code from the Quick start (or define them once), so the multi-turn and common-task snippets are executable.

Link '[training.md](references/training.md)' from the 'Training custom model' section and move the stage details, benchmarks, and performance tables out of SKILL.md, cutting the duplicated quantization content to a single mention.

Add validation/troubleshooting checkpoints (e.g., verifying VRAM fit before loading, what to do on OOM, how to confirm the correct conv template) so the load-and-generate workflow has explicit checkpoints.

DimensionReasoningScore

Conciseness

The body avoids explaining concepts Claude already knows (no 'what is a vision transformer' padding), but contains real duplication: 4-bit quantization is documented twice ('Available models' shows 'load_4bit = True' and a separate 'Quantization (reduce VRAM)' section repeats it), and 'Metrics'/'Benchmarks'/'Performance' tables overlap heavily with the 'When to use' section. Five near-identical 'Common tasks' snippets add little per token. This fits anchor 3 (mostly efficient, some unnecessary content that could be tightened) rather than 4's 'minor instances'.

3 / 5

Actionability

The Quick start and CLI sections are genuinely executable, but several core snippets are pseudocode relying on undefined helpers: the multi-turn example calls an undefined 'generate(conv, model, image)', all five 'Common tasks' call an undefined 'ask(model, image, question)', and the LangChain example returns an undefined 'response'. Anchor 3 ('some concrete guidance but incomplete; pseudocode instead of executable code') fits; it is not 4 because these gaps sit in the skill's primary use cases, and not 2 because the flagship quick-start path is copy-paste runnable.

3 / 5

Workflow Clarity

Sequences exist (clone → install → load → generate; train stage 1 → stage 2) but there are no validation checkpoints or error-recovery guidance anywhere — no troubleshooting for OOM, wrong conv template, or failed installs, and the training stages list commands with no verification steps. This matches anchor 3 (sequence present, checkpoints missing/implicit) rather than 4.

3 / 5

Progressive Disclosure

A bundle file 'references/training.md' exists but is never linked from SKILL.md — the 'Training custom model' section inlines a compressed duplicate of it instead of pointing there, so the reference is undiscoverable. The ~290-line body also inlines material (benchmarks, performance tables, framework integrations) that reads like separate-file content. Anchor 3 ('references present but not clearly signaled; content that should be separate is inline') fits; it is not 2 because the body is well-sectioned with headers, and not 4 because the one existing reference is completely unsignaled.

3 / 5

Total

12

/

20

Passed

Description

53%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description is admirably concise and names a specific niche with three concrete capabilities, but it omits any 'when to use' trigger guidance and lacks natural-language synonyms that would help a user (or Claude) select it over adjacent vision skills. Adding a single 'Use when...' clause would lift its weakest, highest-weighted dimension.

Suggestions

Append an explicit trigger clause, e.g. 'Use when building image chatbots, answering questions about images, or generating captions for pictures.'

Add natural synonyms and phrasings users actually say — 'describe an image', 'visual question answering', 'image chat' — alongside the terse 'VQA'.

Consider naming LLaVA or 'open-source vision-language model' to reduce conflict risk with adjacent skills like CLIP or BLIP-2.

DimensionReasoningScore

Specificity

The description names the domain ('Vision-language chat') and three terse concrete actions ('VQA, captioning, image dialogue'), but coverage is not comprehensive — nothing signals the training, serving, or model-selection capabilities the skill body actually covers. It sits at anchor 3 (domain plus 1-2 concrete actions, not comprehensive) rather than 4, whose example lists several fleshed-out actions with only minor gaps.

3 / 5

Completeness

The 'what' is clear (VQA, captioning, image dialogue) but there is no 'Use when...' clause or any equivalent trigger guidance, which the judging guidelines explicitly cap at 3. It is not a 2 because the 'what' is concrete, not vague, and not a 4 because 'when' is entirely absent rather than merely implicit.

3 / 5

Trigger Term Quality

'VQA', 'captioning', and 'image dialogue' are relevant keywords, but common natural variations users would say ('describe this image', 'image chat', 'visual question answering' spelled out, file extensions) are missing, and 'VQA' leans technical. This matches anchor 3 (some relevant keywords, missing common variations/synonyms) better than anchor 4's 'good keyword coverage with only a few natural terms missing'.

3 / 5

Distinctiveness Conflict Risk

'Vision-language chat' with VQA/captioning/dialogue carves a mostly distinct niche that would only overlap with closely related skills (e.g., a BLIP or CLIP skill), fitting anchor 4. It is not a 5 because the description does not name LLaVA or distinguish itself from other vision-language assistants, so overlap risk with similar multimodal skills remains.

4 / 5

Total

13

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

metadata_field

'metadata' should map string keys to string values

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 2 missing, 2 deeper-than-1-level

Warning

Total

12

/

16

Passed

Repository
NousResearch/hermes-agent
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.