CtrlK
BlogDocsLog inGet started
Tessl Logo

vector-index-tuning

Optimize vector index performance for latency, recall, and memory. Use when tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure.

81

1.56x
Quality

72%

Does it follow best practices?

Impact

100%

1.56x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./tests/ext_conformance/artifacts/agents-wshobson/llm-application-dev/skills/vector-index-tuning/SKILL.md

The canonical home for this skill is vector-index-tuning in wshobson/agents

SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

Technically strong content with concrete, parameter-rich templates and useful reference tables, but it is a monolithic 520-line file that buries executable scripts in the main skill body and lacks an explicit ordered tuning workflow with validation checkpoints. Several templates also contain duplicate or trivial code that inflates token cost.

Suggestions

Move the four code templates into scripts/ or references/ files (e.g., references/benchmarking.md, references/quantization.md) and keep SKILL.md as a concise overview with one-level-deep, clearly signaled links.

Add an explicit ordered workflow — e.g., 1. benchmark current config → 2. estimate memory → 3. select index/quantization → 4. apply → 5. validate recall ≥ target before rollout — with a checkpoint at each step.

Trim duplicated and trivial code (the second calculate_recall in Template 4, the manual bit-packing loop) and fix the small defects: missing `Tuple` import in Template 2 and `tune_search_parameters` returning `SearchParams` against a `dict` annotation.

DimensionReasoningScore

Conciseness

The body is mostly tables and code with little padding of known concepts, but ~400 lines of templates include content Claude can produce itself: a manual bit-packing loop replaceable by `np.packbits`, a second `calculate_recall` implementation duplicating Template 1's, and boilerplate monitor/dataclass plumbing. Not the 4 anchor's 'minor instances of over-explanation' — several templates could be cut or tightened substantially.

3 / 5

Actionability

Four concrete, largely executable templates (hnswlib benchmarking, quantization, Qdrant collection config, monitoring) with real parameter values. Minor gaps keep it from 5: Template 2 uses `Tuple` in signatures without importing it, `tune_search_parameters` is annotated `-> dict` but returns `models.SearchParams`, and `recommend_hnsw_params` ignores `max_latency_ms`/`available_memory_gb`.

4 / 5

Workflow Clarity

There is no explicit tuning workflow — the implied sequence (benchmark → estimate memory → choose quantization → configure → monitor recall) is scattered across templates, and the Best Practices do's/don'ts are unsequenced with no validation checkpoints (e.g., 'verify recall ≥ target before shipping'). The benchmark template does compute recall, which is an implicit checkpoint, so this sits above the 'rough sequence, validation absent' anchor but below 'clear sequence with most checkpoints'.

3 / 5

Progressive Disclosure

Good section structure and headers, but the skill is a 520-line monolith: four full code templates that clearly belong in `scripts/` or `references/` files are inlined, and no bundle files exist or are referenced. Matches the 3 anchor ('some structure... content that should be separate is inline'); not a 2 because the body is well-organized with clear headers rather than a wall of text.

3 / 5

Total

13

/

20

Passed

Description

87%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person voice, concrete capability statement, and an explicit 'Use when' clause with three specific triggers. The only gap is coverage of natural synonyms (embeddings, ANN, faiss, vector database) that would widen trigger matching.

DimensionReasoningScore

Specificity

"Optimize vector index performance for latency, recall, and memory" names the domain and three concrete optimization targets, and the trigger clause lists specific actions ("tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure"). It stops just short of the comprehensive multi-action enumeration of the 5 anchor (e.g., no mention of benchmarking, index-type selection, or monitoring), so it fits the 'several specific actions; minor gaps' anchor.

4 / 5

Completeness

Explicitly answers both: what ("Optimize vector index performance for latency, recall, and memory") and when ("Use when tuning HNSW parameters, selecting quantization strategies, or scaling vector search infrastructure") with concrete trigger phrases. This matches the 5 anchor ('clearly and explicitly answers both what AND when with concrete trigger phrases'); a 4 would require the 'when' to be less explicit or specific, which it is not.

5 / 5

Trigger Term Quality

Strong natural keywords users would say: "HNSW parameters", "quantization", "vector search", "latency", "recall". Missing common synonyms and ecosystem terms a user might naturally use — "embeddings", "nearest neighbor / ANN", "faiss", "qdrant", "vector database" — which the 5 anchor's 'synonyms and file extensions' standard expects.

4 / 5

Distinctiveness Conflict Risk

"HNSW parameters", "quantization strategies", and "vector search infrastructure" carve out a clear niche with distinct technical triggers that no general-purpose or adjacent skill (e.g., generic database tuning) would claim. Minimal conflict risk, matching the 5 anchor.

5 / 5

Total

18

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (524 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
Dicklesworthstone/pi_agent_rust
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.