CtrlK
BlogDocsLog inGet started
Tessl Logo

vision

Query images with a local Ollama vision model without loading the image into the main agent context. Use when you need to describe a screenshot, check whether rendered content is present, detect overlapping elements, or ask any visual question about a PNG/JPEG/WebP file. Requires Ollama running locally with the Gemma 4 multimodal model (`gemma4` on Ollama). Script: .agents/skills/vision/scripts/ask.py. Trigger phrases: "describe image", "what does this screenshot show", "does the canvas contain content", "check screenshot visually", "look at this image", "any overlapping elements", "vision query".

76

Quality

94%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

92%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A strong, highly actionable body: all commands are executable and verified against the bundled script, the workflow is sequenced with a ping checkpoint and a symptom→fix recovery table, and structure is clean with a single one-level script reference. The only issue is minor redundancy — duplicated --info/--memory/--storage listings and slight overlap between Model Selection and Behavior.

Suggestions

Remove the System Info section's duplicated command block (--info/--memory/--storage) and instead reference the Quick Reference, or drop those lines from Quick Reference and keep them only in System Info.

Merge the Model Selection section into Behavior, since Behavior already covers the Gemma-4-only policy and fallback behavior — this would cut ~10 lines of near-duplicate prose.

Note once, next to the Quick Reference, that $SCRIPT must be set before reusing the later examples, instead of relying on the reader having run the first block.

DimensionReasoningScore

Conciseness

The body is lean and assumes Claude's competence (no concept tutorials, no library comparisons), but the --info/--memory/--storage commands appear verbatim in both Quick Reference and System Info, and Model Selection partially restates Behavior's fallback policy. These minor duplications fit the 4 anchor ('minor instances of over-explanation that could be trimmed') rather than 5, where every token earns its place; it is well above 3, which would require genuinely unnecessary explanation.

4 / 5

Actionability

Every example is copy-paste executable (uv run $SCRIPT with real flags: --ping, --info, --memory, --storage, --list-models, --prompt, --model), prerequisites include exact install commands (ollama serve, ollama pull gemma4), and the troubleshooting table maps concrete symptoms to exact fixes — verified against the actual ask.py script, whose flags match. This matches the 5 anchor's 'fully executable, copy-paste ready commands covering the common cases'.

5 / 5

Workflow Clarity

The Typical Agent Workflow gives a clear sequence (image written to disk → targeted query → parse response), the examples explicitly recommend a --ping sanity check before querying, and the Troubleshooting table provides symptom→cause→fix feedback loops for error recovery. The skill involves no destructive or batch operations, so the validation cap does not apply; the single core action is unambiguous, matching the 5 anchor including the simple-skill exception.

5 / 5

Progressive Disclosure

The body is a well-organized overview with clear sections, and its only external pointer is the single script (scripts/ask.py, confirmed to exist in the bundle with matching flags) — one level deep, clearly signaled. Nothing that belongs in a separate file is inlined at ~150 lines of quick-reference material, so navigation is easy, matching the 5 anchor. The 4 anchor's 'minor organization gaps' would require misplaced content, which is absent.

5 / 5

Total

19

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

An excellent description: explicit what and when, concrete third-person actions, and rich natural trigger phrases with file extensions. The only weakness is a couple of overly broad triggers ("look at this image", "ask any visual question") that create minor conflict risk with a main agent's native vision capability.

DimensionReasoningScore

Specificity

The description lists multiple concrete actions — "Query images with a local Ollama vision model", "describe a screenshot", "check whether rendered content is present", "detect overlapping elements", "ask any visual question" — with comprehensive coverage and consistent third-person voice. It matches the 5 anchor (multiple specific concrete actions, comprehensive coverage) rather than 4, whose 'minor gaps in coverage' doesn't apply since querying, describing, verifying, and overlap detection are all named.

5 / 5

Completeness

It explicitly answers both what ("Query images with a local Ollama vision model without loading the image into the main agent context") and when ("Use when you need to describe a screenshot, check whether rendered content is present, detect overlapping elements, or ask any visual question about a PNG/JPEG/WebP file") with concrete trigger phrases, exactly matching the 5 anchor. Score 4 would require the 'when' to be less explicit, which it is not.

5 / 5

Trigger Term Quality

It includes comprehensive natural trigger phrases — "describe image", "what does this screenshot show", "check screenshot visually", "look at this image", "any overlapping elements", "vision query" — plus file extensions "PNG/JPEG/WebP", matching the 5 anchor's requirement for synonyms and extensions. The 4 anchor ('a few natural terms missing') fits worse: common user phrasings for screenshots, content checks, and overlap questions are all present.

5 / 5

Distinctiveness Conflict Risk

The niche is clear — local Ollama vision querying that avoids loading images into the main agent context — with model requirements stated. However, broad triggers like "look at this image" or "ask any visual question" could fire when the main agent's own vision capability or another image skill is the better fit, so the minor overlap risk fits the 4 anchor rather than 5's 'minimal conflict risk'.

4 / 5

Total

19

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
gridaco/grida
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.