CtrlK
BlogDocsLog inGet started
Tessl Logo

pgvector-semantic-search

Use this skill for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search. **Trigger when user asks to:** - Store or search vector embeddings in PostgreSQL - Set up semantic search, similarity search, or nearest neighbor search - Create HNSW or IVFFlat indexes for vectors - Implement RAG (Retrieval Augmented Generation) with PostgreSQL - Optimize pgvector performance, recall, or memory usage - Use binary quantization for large vector datasets **Keywords:** pgvector, embeddings, semantic search, vector similarity, HNSW, IVFFlat, halfvec, cosine distance, nearest neighbor, RAG, LLM, AI search Covers: halfvec storage, HNSW index configuration (m, ef_construction, ef_search), quantization strategies, filtered search, bulk loading, and performance tuning.

71

Quality

89%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide
SecuritybySnyk

Passed

No findings from the security scan

SKILL.md
Quality
Evals
Security

Quality

Content

78%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

An unusually strong single-file reference: executable SQL everywhere, concrete tuning values, honest uncertainty labels on benchmarks, and a symptom→fix troubleshooting table that serves as an error-recovery feedback loop. The main gaps are structural — it is a ~320-line monolith with no progressive disclosure via reference files — plus a slightly padded conceptual opening and no consolidated end-to-end setup→validation sequence.

Suggestions

Move detail-heavy material (the RAM/capacity tables, binary quantization deep-dive, pgvectorscale alternative, and bulk-loading specifics) into a references/ file such as references/quantization.md, keeping SKILL.md as a lean overview with clearly signaled one-level-deep links.

Replace the opening conceptual paragraph about what semantic search and embedding models are with a single sentence, since this is knowledge Claude already has.

Consolidate the Golden Path, Core Rules, and Performance by Dataset Size sections into one ordered setup→load→index→query→validate sequence with explicit validation checkpoints (EXPLAIN ANALYZE and recall-comparison) at each stage.

DimensionReasoningScore

Conciseness

The body is dense and efficient — nearly every section delivers pgvector-specific tuning knowledge Claude cannot be assumed to know (ef_search recall/speed tables, halfvec capacity-per-RAM tables, iterative-scan budgets, an 80x oversampling rationale), with no library-selection fluff. It drops to 4 rather than 5 because the opening paragraph explains concepts Claude already knows ('Semantic search finds content by meaning rather than exact keywords. An embedding model converts text into high-dimensional vectors, where similar meanings map to nearby points') and some table rows restate the same recall/latency tradeoff three times (ef_search table, 'Performance by Dataset Size', and 'Key rules').

4 / 5

Actionability

Guidance is fully executable and copy-paste ready throughout: complete CREATE TABLE / CREATE INDEX statements, exact parameter values ('m = 16, ef_construction = 64', 'SET hnsw.ef_search = 100', 'lists = 1000'), a complete working binary-quantization query with re-ranking, bulk-loading COPY and maintenance_work_mem commands, and debugging SQL (EXPLAIN (ANALYZE, BUFFERS), enable_indexscan toggles). This matches the score-5 anchor ('copy-paste ready code or commands; specific examples cover the common cases'); even edge conditions like operator/index-ops mismatch and dimension casting are handled concretely.

5 / 5

Workflow Clarity

The 'Golden Path (Default Setup)' gives an unambiguous default sequence (data type → distance → index → ef_search → query pattern), ordering rules are explicit ('Index after bulk loading', 'Create indexes concurrently in production'), and the 'Common Issues (Symptom → Fix)' table plus recall-comparison SQL provide real validation and error-recovery loops for the batch bulk-load operation. It falls short of 5 because there is no end-to-end numbered sequence connecting setup → load → index → query → validate; validation steps are scattered across Monitoring & Debugging and prose notes ('Validate by monitoring cache residency and p95/p99 latency') rather than presented as explicit checkpoints in the main workflow.

4 / 5

Progressive Disclosure

The file is well-organized with clear section headers and the only external pointers (pgvector README, pgvectorscale) are clearly signaled, but it is a ~320-line monolith with no bundle files at all — content that naturally belongs in separate references, such as the RAM-capacity tables, the full binary-quantization deep dive, and the pgvectorscale alternative, is inlined. This matches the score-3 anchor ('Some structure but could be better organized; ... content that should be separate is inline'); it is above score 2 because headers, tables, and a symptom→fix section make it genuinely navigable, but below score 4 because a skill this long would benefit from splitting detail into reference files.

3 / 5

Total

16

/

20

Passed

Description

96%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: it states concrete capabilities, explicitly enumerates trigger conditions, and provides a rich keyword list with synonyms. Its only weaknesses are minor — the keyword block borders on padded ('LLM', 'AI search') and slightly raises conflict risk with general RAG/embedding skills, and the multi-paragraph structure is longer than the ideal one-to-two sentence description.

DimensionReasoningScore

Specificity

The description lists multiple concrete, specific actions — 'Store or search vector embeddings in PostgreSQL', 'Create HNSW or IVFFlat indexes for vectors', 'Implement RAG (Retrieval Augmented Generation) with PostgreSQL', 'Use binary quantization for large vector datasets' — plus a 'Covers:' line enumerating 'halfvec storage, HNSW index configuration (m, ef_construction, ef_search), quantization strategies, filtered search, bulk loading, and performance tuning'. This matches the score-5 anchor ('Lists multiple specific concrete actions; comprehensive coverage'); score 4 would require minor coverage gaps, but even advanced topics (filtered search, bulk loading) are explicitly named.

5 / 5

Completeness

It explicitly answers both questions: 'what' is stated upfront ('Use this skill for setting up vector similarity search with pgvector for AI/ML embeddings, RAG applications, or semantic search') and 'when' is given as an explicit trigger block ('**Trigger when user asks to:**' with six concrete conditions). This matches the score-5 anchor ('Clearly and explicitly answers both what AND when with concrete trigger phrases'); the explicit trigger guidance means the completeness cap of 3 for a missing 'Use when' clause does not apply.

5 / 5

Trigger Term Quality

Keyword coverage is comprehensive and includes natural user phrasings and synonyms: 'pgvector, embeddings, semantic search, vector similarity, HNSW, IVFFlat, halfvec, cosine distance, nearest neighbor, RAG, LLM, AI search', alongside trigger phrases users would naturally say ('Set up semantic search, similarity search, or nearest neighbor search', 'Optimize pgvector performance, recall, or memory usage'). This matches the score-5 anchor (comprehensive coverage of natural terms including synonyms); a term like 'vector database' is missing but coverage clearly exceeds the score-4 anchor's 'a few natural terms missing'.

5 / 5

Distinctiveness Conflict Risk

The skill has a clear niche — pgvector/PostgreSQL vector search — with distinctive terms (pgvector, halfvec, HNSW, IVFFlat) that would rarely collide with unrelated skills, matching 'Mostly distinct; minor overlap risk'. However, broad keywords like 'RAG, LLM, AI search' and 'embeddings' overlap with general RAG/embedding-model or other vector-database skills, so it falls below the score-5 anchor's 'minimal conflict risk'.

4 / 5

Total

19

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

Total

15

/

16

Passed

Repository
timescale/pg-aiguide
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.