CtrlK
BlogDocsLog inGet started
Tessl Logo

llava

Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.

56

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/llava/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is dense with executable code, commands, and useful comparison tables, but it is over-long for a SKILL.md: it duplicates quantization and VRAM content, pads the 'Common tasks' section with five near-identical pseudocode snippets using an undefined helper, and never points to the existing references/training.md file. Tightening the inline duplication and linking the training reference would lift the weakest dimensions.

Suggestions

Replace the five near-identical 'Common tasks' snippets (which call an undefined ask() helper) with a short table of example prompts, or define the ask() helper once — this addresses both the conciseness and actionability gaps.

Link the 'Training custom model' section to the existing references/training.md (e.g., 'See [training.md](references/training.md) for stage configuration, data, and hardware details') and remove the duplicated VRAM/quantization content between 'Available models', 'Quantization', and 'Performance'.

Add validation checkpoints to the workflows: after the two training stages, note how to verify the checkpoint (e.g., run a sample VQA prompt and check output sanity), and after Quick start inference, suggest a quick sanity check of the decoded response.

DimensionReasoningScore

Conciseness

Mostly efficient, concrete reference material (installation, model loading, CLI, tables), but padded sections remain: five near-identical 'Common tasks' snippets, quantization guidance duplicated in 'Available models' and its own 'Quantization' section, VRAM figures repeated in two tables, and marketing-style metrics ('23,000+ GitHub stars', 'GPT-4V level capabilities (targeted)'). This fits anchor 3's 'some unnecessary explanation or could be tightened' better than anchor 2, since the padding is a minority of the body.

3 / 5

Actionability

The Quick start is complete and executable (model loading, image processing, generation), and CLI/Gradio/quantization commands are concrete. It misses anchor 5 because several secondary examples rely on undefined helpers — 'response = ask(model, image, question)', 'generate(conv, model, image)', and the LangChain stub returning an undefined 'response' — so common cases are not uniformly copy-paste ready.

4 / 5

Workflow Clarity

Sequences exist (training Stage 1 pretrain.sh → Stage 2 finetune.sh; multi-turn conversation turns), but no validation or verification checkpoints appear anywhere — no output checks after training, no sanity check after inference, and the multi-turn example's step 'conv.messages[-1][1] = response1' is left implicit. This matches anchor 3 ('steps listed but validation gaps') rather than anchor 4.

3 / 5

Progressive Disclosure

The bundle contains references/training.md (a 197-line training guide), but the body never links to it — the 'Training custom model' section inlines a two-line duplicate instead of signaling the reference. Section headers within the ~290-line body are clear, but the un-signaled reference file and inlined training content match anchor 3 ('references present but not clearly signaled; content that should be separate is inline').

3 / 5

Total

13

/

20

Passed

Description

75%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid description: specific, third-person, with an explicit 'Use for...' trigger clause and mostly natural terminology. Its main weaknesses are trigger phrasing that describes task categories rather than concrete user utterances, and missing common synonyms like VQA or image captioning.

DimensionReasoningScore

Specificity

The description lists several concrete capabilities — 'Enables visual instruction tuning and image-based conversations', 'Supports multi-turn image chat, visual question answering, and instruction following' — in third-person voice. It sits below anchor 5 because coverage has minor gaps (image captioning/description and document understanding from the body are absent) and above anchor 3 because more than 1-2 specific actions are named.

4 / 5

Completeness

Both 'what' (visual instruction tuning, multi-turn image chat, VQA, instruction following, CLIP+Vicuna/LLaMA architecture) and an explicit 'when' ('Use for vision-language chatbots or image understanding tasks') are present. It falls short of anchor 5 because the trigger clause is task-category-oriented rather than phrased as concrete user triggers ('Use when the user asks about an image...'), and exceeds anchor 3 because the 'when' is explicit, not merely implied.

4 / 5

Trigger Term Quality

Natural phrases a user would say are present: 'image chat', 'visual question answering', 'image understanding', 'vision-language chatbots'. It is below anchor 5 because common variations like 'VQA', 'image captioning', or 'describe an image' are missing, and above anchor 3 because multiple natural terms beyond generic jargon are covered.

4 / 5

Distinctiveness Conflict Risk

The niche — conversational image analysis with a CLIP+Vicuna/LLaMA architecture, 'Best for conversational image analysis' — is distinct from zero-shot classifiers or captioning-only tools. Minor overlap risk remains with other vision-language model skills (e.g., BLIP-2-style skills), keeping it below anchor 5.

4 / 5

Total

16

/

20

Passed

Validation

81%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 13 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

referenced_paths_exist

Referenced path issues: 2 missing, 2 deeper-than-1-level

Warning

Total

13

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.