CtrlK
BlogDocsLog inGet started
Tessl Logo

chroma

Open-source embedding database for AI applications. Store embeddings and metadata, perform vector and full-text search, filter by metadata. Simple 4-function API. Scales from notebooks to production clusters. Use for semantic search, RAG applications, or document retrieval. Best for local development and open-source projects.

60

Quality

70%

Does it follow best practices?

Run evals on this skill

Adds up to 20 points to the overall score

View guide

SecuritybySnyk

Passed

No findings from the security scan

Fix and improve this skill with Tessl

tessl review fix ./backend/cli/skills/llm-tools/chroma/SKILL.md
SKILL.md
Quality
Evals
Security

Quality

Content

57%Weight 40%Scale 1-5

Reviews the quality of instructions and guidance provided to agents. Good implementation is clear, handles edge cases, and produces reliable results.

A well-organized, largely executable API reference for Chroma. Its weaknesses are padding (marketing metrics, generic best-practices, duplicated examples), no validation guidance around destructive delete operations, and failure to use the provided references/integration.md — whose integration content is inlined instead of referenced.

Suggestions

Replace the inlined LangChain/LlamaIndex and Resources sections with pointers to the existing references/integration.md (e.g., "**LangChain/LlamaIndex**: See [integration.md](references/integration.md)"), removing the duplication and shrinking the body.

Add a validate-before-delete checkpoint: show `collection.get(where=...)` to preview matched documents before `collection.delete(where=...)` or `client.delete_collection(...)`, and caution that deletes are irreversible.

Cut the tagline, GitHub-stars/forks metrics, Discord link, and latency table, and fold the 10 generic best-practice bullets down to the few Chroma-specific ones (persistent client, metadata on add, batch adds).

DimensionReasoningScore

Conciseness

The bulk is lean code, but there is more than minor padding: a marketing tagline ("The AI-native database for building LLM applications with memory"), a GitHub-stars/forks metrics block, a Discord link, a 10-item best-practices list of generic advice ("Unique IDs - Avoid collisions"), a rough latency table, time-sensitive version data ("v1.3.3 (stable, weekly releases)") outside any old-patterns section, and the OpenAI embedding-function example repeated verbatim. This matches the anchor for mostly efficient with some unnecessary content; it is not 4 because the duplication and promotional sections are more than trimmable instances, and not 2 because the core API content is tight.

3 / 5

Actionability

Mostly executable, copy-paste-ready Python for installation, CRUD, persistence, filtering, and server mode. Not 5: the custom embedding function example is pseudocode (returns an undefined `embeddings` variable), the custom-embeddings example uses literal ellipses ("[[0.1, 0.2, ...]]"), and the LangChain example imports `RecursiveCharacterTextSplitter` from `langchain` rather than `langchain_text_splitters`. Not 3: these are minor gaps in otherwise concrete, runnable guidance.

4 / 5

Workflow Clarity

The operations are organized into unambiguous single-purpose sections, but the skill documents destructive batch operations — `client.delete_collection("my_docs")` and `collection.delete(where={"source": "outdated"})` — with no validation or verification guidance (e.g., previewing the matched documents with `collection.get(where=...)` before deleting). Per the rubric, missing validation on destructive/batch operations caps this at 3, taking precedence over the simple-skill exception; it is not 2 because each operation is clearly and coherently presented.

3 / 5

Progressive Disclosure

Sections are well-structured with clear headers, but the bundle provides `references/integration.md` and the body never signals or links to it — instead the LangChain, LlamaIndex, and Resources content is fully inlined and duplicated there. This matches the anchor for structure present but references not clearly signaled and separable content inline. Not 4: the one reference file that exists is entirely unlinked and its content is redundantly inlined; not 2: the document itself has genuine, navigable section structure.

3 / 5

Total

13

/

20

Passed

Description

83%Weight 40%Scale 1-5

Based on the skill's description, can an agent find and select it at the right time? Clear, specific descriptions lead to better discovery.

A strong description: third-person, concrete about capabilities, and explicit about when to use it with real trigger phrases. Its main weaknesses are missing synonyms ("vector database", "similarity search", "ChromaDB") and slightly incomplete action coverage relative to the body's content.

DimensionReasoningScore

Specificity

"Store embeddings and metadata, perform vector and full-text search, filter by metadata" lists several concrete actions, matching the anchor for several specific actions with minor gaps. Not 5: coverage is not comprehensive — update/delete operations, persistence, and server mode are absent, and "Simple 4-function API" / "Scales from notebooks to production clusters" are marketing framing rather than concrete actions. Not 3: more than 1-2 concrete actions are explicitly listed.

4 / 5

Completeness

What is explicit and concrete ("Store embeddings and metadata, perform vector and full-text search, filter by metadata") and when is explicit with concrete trigger phrases ("Use for semantic search, RAG applications, or document retrieval. Best for local development and open-source projects"), matching the anchor that clearly answers both with concrete triggers. Not 4: the when-clause is more specific than the anchor-4 example's generic "Use when working with X", listing three distinct application types plus deployment context.

5 / 5

Trigger Term Quality

Natural phrases users would say are present: "semantic search", "RAG applications", "document retrieval", "vector and full-text search", "embeddings". Not 5: common variations and synonyms are missing — no "vector database", "similarity search", "Chroma/ChromaDB", or "vector store" wording. Not 3: keyword coverage goes well beyond a single generic term.

4 / 5

Distinctiveness Conflict Risk

The embedding-database niche is clearly staked out and "Best for local development and open-source projects" distinguishes it from managed alternatives. Not 5: the description itself never names Chroma/ChromaDB, and in a RAG skill collection it overlaps with other vector-database skills (Pinecone, FAISS, Qdrant). Not 3: it is more specific than a generic "works with vector databases" skill.

4 / 5

Total

17

/

20

Passed

Validation

87%

Checks the skill against the spec for correct structure and formatting. All validation checks must pass before discovery and implementation can be scored.

Validation — 14 / 16 Passed

Validation for skill structure

CriteriaDescriptionResult

metadata_version

'metadata.version' is missing

Warning

frontmatter_unknown_keys

Unknown frontmatter key(s) found; consider removing or moving to metadata

Warning

Total

14

/

16

Passed

Repository
synthetic-sciences/openscience
Reviewed

Table of Contents

Is this your skill?

If you maintain this skill, you can claim it as your own. Once claimed, you can manage eval scenarios, bundle related skills, attach documentation or rules, and ensure cross-agent compatibility.