CtrlK
BlogDocsLog inGet started
Tessl Logo

clip

OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.

64

Quality

76%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/clip/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

65%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized, code-heavy reference with mostly executable examples, though a couple (the vector-database integration) are not self-contained. Its main structural flaws are the orphaned references/applications.md that SKILL.md never points to (with heavily duplicated application content inlined) and the absence of any validation or feedback loop despite covering batch and moderation operations, capping workflow clarity at 3.

Suggestions

Link the existing references/applications.md from the body (e.g. a '## Applications' section pointing to it) and move the duplicated inlined recipes — content moderation, semantic image search, batch processing, best practices — into that file, keeping only the quick-start and model-selection guidance in SKILL.md.

Add validation/verification steps to the batch and content-moderation workflows (e.g. verify embedding shapes before database insertion, or check confidence thresholds before acting on a moderation verdict) so workflow clarity can exceed the cap of 3.

Fix the chromadb integration example to be self-contained: move tokenized text to device, wrap encoding in torch.no_grad(), and define image_paths/image_embeddings locally instead of depending on an earlier section.

DimensionReasoningScore

Conciseness

The body is dominated by dense, practical code with minimal conceptual explanation, but a few sections could be trimmed: the "Metrics" section ("25,300+ GitHub stars") is promotional padding, the opening sentence repeats the frontmatter description, and the labels list is defined twice in the quick-start example. Mostly lean with minor over-explanation, fitting anchor 4; not 5 because of the metrics/stars padding and repeated boilerplate.

4 / 5

Actionability

Most sections give copy-paste-ready executable code (zero-shot classification, similarity, semantic search, batch processing), but minor gaps exist: the chromadb snippet calls model.encode_text(clip.tokenize([query])) without moving tokens to device or wrapping in no_grad, and it reuses variables (image_embeddings, image_paths) from a prior section, so it won't run standalone. Anchor 4 ("mostly executable... minor gaps") fits better than 5 ("copy-paste ready, covers common cases") because the integration example is not self-contained.

4 / 5

Workflow Clarity

Content is organized as parallel recipes rather than a sequenced workflow, with no validation or verification checkpoints anywhere — e.g. the "Batch processing" and "Content moderation" sections never verify results or handle failures. Per the rubric, a skill with batch operations and no validation/feedback loop is capped at 3, which takes precedence; the sequence that exists (install → load → encode → compare) is coherent but checkpoints are absent.

3 / 5

Progressive Disclosure

The body has clear section structure, but the bundle's references/applications.md is never linked or signaled from SKILL.md, and the body inlines application recipes (zero-shot classification, semantic image search, content moderation, best practices) that duplicate that file's content. Anchor 3 ("references present but not clearly signaled; content that should be separate is inline") matches; not 4 because a provided bundle file is completely orphaned and substantial duplicated content is inlined.

3 / 5

Total

14

/

20

Passed

Description

88%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it names the tool, lists concrete capabilities in third person, and includes an explicit "Use for..." trigger clause with natural keywords. The only weakness is moderate overlap risk with other vision-language skills and missing a few natural trigger variations.

DimensionReasoningScore

Specificity

"Enables zero-shot image classification, image-text matching, and cross-modal retrieval" plus "Use for image search, content moderation, or vision-language tasks" lists multiple specific concrete actions with comprehensive coverage of the model's capabilities. Voice is third person throughout ("OpenAI's model connecting vision and language"), so no person-voice penalty applies.

5 / 5

Completeness

Explicitly answers both: what ("Enables zero-shot image classification, image-text matching, and cross-modal retrieval") and when ("Use for image search, content moderation, or vision-language tasks without fine-tuning") with concrete trigger phrases. The "Use for..." clause satisfies the explicit-trigger-guidance requirement, so the score-3 cap does not apply.

5 / 5

Trigger Term Quality

Good natural keyword coverage — "image search", "content moderation", "image classification", "vision-language tasks" — but misses common variations users would say such as "photo search", "NSFW detection", "image similarity", or "matching an image to text". Not score 3 because several relevant natural phrases are present; not score 5 because synonyms and variations are incompletely covered.

4 / 5

Distinctiveness Conflict Risk

"Without fine-tuning" and the zero-shot framing carve a clear CLIP niche, but "vision-language tasks" and "image search" broadly overlap with closely related skills (BLIP-2, LLaVA, other multimodal retrieval tools). Mostly distinct with minor overlap risk against closely related vision-language skills, fitting anchor 4 rather than the minimal-conflict anchor 5.

4 / 5

Total

18

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.