CtrlK
BlogDocsLog inGet started
Tessl Logo

add-image-vision

Add image vision to NanoClaw agents. Resizes and processes WhatsApp image attachments, then sends them to Claude as multimodal content blocks.

68

Quality

85%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Low

Low-risk findings worth noting

SKILL.md
Quality
Evals
Security

Quality

Content

96%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body is an exemplary phased runbook: concise, fully executable, with explicit validation gates and error-recovery troubleshooting. Its only structural limitation is that it is a self-contained single file with no progressive disclosure to detailed material, which is largely appropriate given the skill's scope.

DimensionReasoningScore

Conciseness

The body is a lean operational runbook: every phase (pre-flight, apply, validate, configure, verify, troubleshooting) carries commands or checks, with no explanation of concepts Claude already knows (no git tutorial, no "what is sharp" padding). No token fails to earn its place, so it is not 4.

5 / 5

Actionability

Fully executable, copy-paste-ready commands throughout: "git remote add whatsapp ...", "git fetch whatsapp skill/image-vision", the merge snippet with a package-lock.json conflict fallback, "npm run build", "npx vitest run src/image.test.ts", "./container/build.sh", the cp sync loop, "launchctl kickstart -k ...", and "tail -50 groups/*/logs/container-*.log". The conflict fallback covers the common failure case, matching the top anchor.

5 / 5

Workflow Clarity

A clearly sequenced multi-step process with explicit validation checkpoints and feedback loops: a pre-flight existence check, the gate "All tests must pass and build must be clean before proceeding", a dedicated verify phase with log greps, and troubleshooting entries that pair error messages with recovery actions (e.g. "Image - processing failed" → "npm ls sharp"). This matches the validate→fix→retry anchor; not 4 because checkpoints are explicit rather than implicit.

5 / 5

Progressive Disclosure

Sections are well organized and nothing in the body clearly belongs in a separate file (no bulk API reference or reference material is inlined), and no bundle files exist to navigate. Not 5 because the body is ~89 lines (over the 50-line simple-skill exception) and there is no reference navigation to signal; not 3 because structure and content placement are good, not merely "could be better organized".

4 / 5

Total

19

/

20

Passed

Description

66%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

The description states concrete, specific capabilities in third person with a distinct niche (NanoClaw + WhatsApp images), but it lacks any "when to use" trigger guidance, which caps its completeness and slightly weakens trigger quality and distinctiveness. Adding an explicit "Use when..." clause with natural synonyms (photos, .jpg/.png, see images) would lift it to top marks.

Suggestions

Append an explicit trigger clause, e.g. "Use when a NanoClaw agent needs to see or understand images sent via WhatsApp (photos, screenshots, .jpg/.png attachments)."

Replace or qualify the generic verb "processes" with the concrete operations the skill actually performs (e.g., "downloads, resizes, and base64-encodes").

Include common user synonyms such as "photos" and file extensions like ".jpg/.png" alongside "image attachments" to improve trigger term coverage.

DimensionReasoningScore

Specificity

"Resizes and processes WhatsApp image attachments, then sends them to Claude as multimodal content blocks" lists several concrete actions (resize, process, send as multimodal blocks), but "processes" is generic and coverage has minor gaps. Not 5 because coverage is not comprehensive; not 3 because it goes beyond naming just 1-2 actions.

4 / 5

Completeness

The "what" is clear ("Add image vision to NanoClaw agents. Resizes and processes WhatsApp image attachments, then sends them to Claude as multimodal content blocks") but there is no "Use when..." clause; the "when" is only weakly implied via the WhatsApp mention. Per the rubric guideline, a missing "Use when..." clause caps completeness at 3.

3 / 5

Trigger Term Quality

Natural terms users would say are present — "image", "vision", "WhatsApp", "attachments", "multimodal" — giving good keyword coverage. Not 5 because common synonyms and extensions are missing ("photos", "see images", ".jpg/.png"); not 3 because coverage is good rather than merely "some relevant keywords".

4 / 5

Distinctiveness Conflict Risk

A clear niche is staked out by "NanoClaw agents" + "WhatsApp image attachments" + "multimodal content blocks", leaving only minor overlap risk with other WhatsApp-channel skills. Not 5 because the absent trigger clause makes it less distinct than the anchor's "clear niche with distinct triggers".

4 / 5

Total

15

/

20

Passed

Validation

100%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 16 / 16 Passed

Validation for skill structure

No warnings or errors.

Repository
jbaruch/nanoclaw-telegram
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.