CtrlK
BlogDocsLog inGet started
Tessl Logo

hybrid-search-implementation

Combine vector and keyword search for improved retrieval. Use when implementing RAG systems, building search engines, or when neither approach alone provides sufficient recall.

74

1.13x
Quality

63%

Does it follow best practices?

Impact

93%

1.13x

Average score across 3 eval scenarios

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./tests/ext_conformance/artifacts/agents-wshobson/llm-application-dev/skills/hybrid-search-implementation/SKILL.md

The canonical home for this skill is hybrid-search-implementation in wshobson/agents

SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

The body delivers genuinely actionable, executable templates for the main hybrid-search backends, with lean prose and good structure. Its weaknesses are redundancy (fusion implemented three ways inline), the absence of any validation or tuning feedback loops, and no progressive disclosure - everything lives in one large file with no reference bundle.

Suggestions

Move the backend-specific templates (PostgreSQL, Elasticsearch, RAG pipeline) into references/ files and keep a concise fusion-method overview plus one minimal example in SKILL.md.

Fix the executable gaps: add `import asyncio` to Template 4 and clarify Template 2's positional SQL parameters ($1-$4).

Add a concrete validation loop for weight tuning, e.g., build a small labeled query set, grid-search alpha/RRF k, and compare recall@k before adopting weights.

DimensionReasoningScore

Conciseness

The prose is lean, but roughly 500 lines of code are inlined and fusion logic is implemented three times (Template 1's RRF function, the SQL RRF in Template 2, and _rrf_fusion in Template 4), so the body could be tightened considerably.

3 / 5

Actionability

The templates are concrete, executable Python/SQL with docstrings and cover the common backends (PostgreSQL/pgvector, Elasticsearch, generic pipeline). Minor gaps keep it below fully copy-paste ready: Template 4 calls asyncio.gather without importing asyncio, and Template 2's SQL parameter numbering ($1-$4) is hard to follow.

4 / 5

Workflow Clarity

The architecture diagram and Template 4's Step 1-4 sequence give a rough flow, but there are no validation checkpoints or feedback loops; directives like "Tune weights empirically - Test on your data" and "A/B test" give no how. This is not a destructive/batch skill, so the cap is not triggered, but checkpoints are missing.

3 / 5

Progressive Disclosure

No bundle files exist (no references/, scripts/, or assets/), and all four full backend-specific templates are inlined in a single ~570-line SKILL.md. Section headers and the fusion comparison table provide real structure, lifting it above the no-structure anchor, but content that clearly belongs in separate reference files is inline.

3 / 5

Total

13

/

20

Passed

Description

70%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A solid description with a clear what, an explicit 'Use when' clause, and good natural trigger terms. It falls short of the top anchors on specificity (a single action rather than a list of capabilities) and on trigger synonyms, and the recall-based trigger is abstract rather than user-voiced.

Suggestions

List 2-3 more concrete capabilities (e.g., 'Fuse result lists with RRF or weighted linear combination, rerank with cross-encoders') to raise specificity.

Add natural synonyms users would say, such as 'hybrid search', 'semantic search', or 'BM25', to the trigger clause.

Replace the abstract trigger 'when neither approach alone provides sufficient recall' with user-voiced conditions like 'when users report missing exact matches for names, codes, or IDs'.

DimensionReasoningScore

Specificity

Names the domain and one concrete action ("Combine vector and keyword search for improved retrieval") but does not enumerate several specific capabilities, matching the anchor for 1-2 concrete actions without comprehensive coverage.

3 / 5

Completeness

Both what ("Combine vector and keyword search for improved retrieval") and when ("Use when implementing RAG systems, building search engines...") are present; the third trigger "when neither approach alone provides sufficient recall" is abstract rather than a concrete phrase a user would naturally say.

4 / 5

Trigger Term Quality

Includes natural user phrases like "RAG systems", "building search engines", "vector and keyword search", and "recall", but misses common synonyms such as "hybrid search", "semantic search", "BM25", or "embeddings".

4 / 5

Distinctiveness Conflict Risk

The RAG/search-engine framing carves a fairly distinct niche with explicit triggers, with only minor overlap risk against generic RAG or pure vector-search skills.

4 / 5

Total

15

/

20

Passed

Validation

93%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 15 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

skill_md_line_count

SKILL.md is long (571 lines); consider splitting into references/ and linking

Warning

Total

15

/

16

Passed

Repository
Dicklesworthstone/pi_agent_rust
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.