CtrlK
BlogDocsLog inGet started
Tessl Logo

blip-2-vision-language

Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.

56

Quality

66%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/blip-2/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

50%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is a competent, code-dense reference: nearly every section gives runnable transformers/LAVIS examples with real model IDs and a useful variants/VRAM table. Its main flaws are length and redundancy — ~570 lines that re-teach known concepts and inline three full workflow classes plus advanced optimization content that belongs in references/advanced-usage.md — and the absence of any validation checkpoints in its batch-processing workflows.

Suggestions

Move the three full workflow classes (ImageCaptioner, VisualQA, ImageSearchEngine), the quantization and Flash Attention sections, and the LAVIS ITM/feature-extraction snippets into references/advanced-usage.md, leaving SKILL.md as a lean quick-start with the model-variant and VRAM tables plus the reference links.

Cut concept-teaching content Claude already knows: the ASCII architecture diagram, the Q-Former components table, and inline comments re-explaining standard generate() parameters (num_beams, top_p, temperature).

Add validation checkpoints to the batch workflows — e.g., verify decoded captions are non-empty and shapes of extracted features before indexing — so batch operations include a verify-and-retry loop.

DimensionReasoningScore

Conciseness

At ~570 lines the body teaches concepts Claude already knows: an ASCII architecture diagram of Q-Former, a "Q-Former components" parameter table, and generation-parameter comments like "top_p=0.9, # Nucleus sampling" and "temperature=0.7, # Creativity" re-explain standard transformers settings. Several sections also restate the same load-and-generate flow (transformers quick start, LAVIS quick start, three full workflow classes). This is noticeably verbose with several padded sections (anchor 2), though the bad_overall-style conceptual essay is avoided, keeping it above a 1.

2 / 5

Actionability

Most examples are concrete, executable code with real model IDs ("Salesforce/blip2-opt-2.7b", BitsAndBytesConfig 8-bit/4-bit snippets, complete ImageCaptioner/VisualQA classes). Not a 5: several snippets depend on undefined variables — "raw_image" and "device" in the ITM and feature-extraction sections, and "processor"/"inputs" in "Controlling generation" — so they are not copy-paste ready without stitching in earlier context. Comfortably above a 3, which would mean pseudocode or missing key details.

4 / 5

Workflow Clarity

The quick start (install → caption → VQA) is a clear sequence, but the skill contains batch operations ("Batch processing", "caption_batch", "index_images" over many files) with no validation or verification checkpoints — e.g., no check that captions decoded correctly or that indexed features are non-degenerate. Per the rubric, missing validation in batch workflows caps this dimension at 3: sequence present, checkpoints absent. It is not a 2, because steps are individually well-defined with concrete commands.

3 / 5

Progressive Disclosure

The body does link two real, clearly-labeled reference files ("[Advanced Usage](references/advanced-usage.md)", "[Troubleshooting](references/troubleshooting.md)", both present in the bundle), one level deep. However, substantial advanced content is inlined in SKILL.md instead of those files — quantization, Flash Attention, feature extraction, and three full ~50-line workflow classes — which is exactly the anchor-3 pattern of content that should be separate sitting inline. Not a 4: more than a minor amount of misplaced content; not a 2, because structure and reference signaling are genuine.

3 / 5

Total

12

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states what the skill does in one crisp sentence and gives an explicit 'Use when' clause with four concrete, natural trigger tasks. Its weaknesses are promotional fluff ("state-of-the-art zero-shot performance") and missing common synonyms/acronyms (VQA, BLIP-2) that would improve trigger matching.

DimensionReasoningScore

Specificity

"Vision-language pre-training framework bridging frozen image encoders and LLMs" names the domain, and the description lists several concrete actions ("image captioning, visual question answering, image-text retrieval, or multimodal chat"). Not a 5: "state-of-the-art zero-shot performance" is promotional padding rather than a concrete capability, and fine-tuning/deployment tasks covered in the body are absent. Clearly above a 3, which expects only 1-2 concrete actions.

4 / 5

Completeness

The description explicitly answers both questions: what ("Vision-language pre-training framework bridging frozen image encoders and LLMs") and when ("Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat") with concrete trigger phrases. This matches the score-5 good example's structure; the score-4 anchor would require a vaguer 'when' clause than the enumerated trigger list given here.

5 / 5

Trigger Term Quality

Triggers include natural phrases users would say: "image captioning", "visual question answering", "image-text retrieval", "multimodal chat". Not a 5: misses common variations users actually type — the acronym "VQA", "BLIP-2" itself, "describe an image", "image-text similarity" — so a few natural terms are missing, matching the score-4 anchor.

4 / 5

Distinctiveness Conflict Risk

"Vision-language pre-training framework" with triggers like image captioning and image-text retrieval carves out a mostly distinct niche. Not a 5: the vision-language space has closely related skills (CLIP for retrieval/similarity, LLaVA/InstructBLIP for multimodal chat) that share trigger terms, creating minor overlap risk — exactly the score-4 anchor.

4 / 5

Total

17

/

20

Passed

Validation

75%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 12 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (574 lines); consider splitting into references/ and linking

Warning

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

relative_links

Relative link issues: 7 missing

Warning

Total

12

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.