Build Semantic Search with ChromaDB in Python (2026)
Build Semantic Search with ChromaDB in Python (2026) shows how to store documents, auto-embed them with the default MiniLM model, and run cosine similarity queries for RAG — all in-process with PersistentClient. No API keys, no separate vector server for local prototypes.
Pair Chroma with Instructor for structured LLM outputs when you turn retrieved chunks into typed answers, or with DSPy to optimize LLM pipelines around your retriever. For PDF-to-text ingest before you index, see Docling PDF to Markdown.
TL;DR
- ChromaDB is an embedded vector store:
add()documents,query()by meaning, filter withwhere. - Install with
pip install chromadb(we tested1.5.9on Python 3.13.5). - Use
PersistentClient(path=...)so collections survive process restarts. - Default embedding is
all-MiniLM-L6-v2(downloaded once into the local cache). - Real run: RAG query ranked the ChromaDB snippet first at distance
0.2465.
Why ChromaDB in 2026?
Most RAG tutorials need a place to put embeddings. Chroma remains the default “pip install and go” choice: one Python package, local persistence, metadata filters, and a tiny API surface. Hosted Chroma Cloud exists when you outgrow a laptop disk, but the local client is enough to learn the patterns and ship a prototype.
| Store | Best for | Local start |
|---|---|---|
| ChromaDB | RAG prototypes, notebooks, small apps | pip install chromadb |
| LanceDB | Hybrid / columnar vector tables | Separate package + hybrid APIs |
| Qdrant / pgvector | Dedicated server or Postgres | Extra process or DB |
Versions tested (2026-09-30)
- Python
3.13.5 chromadb1.5.9- Default embedding:
all-MiniLM-L6-v2(ONNX, auto-downloaded) - Demo path:
./chroma_dataviaPersistentClient
python -m venv .venv && source .venv/bin/activate
pip install chromadb==1.5.9
python rag_demo.py
1. PersistentClient + collection
Save this as rag_demo.py. We wipe a fresh chroma_data folder so the run is reproducible, create a cosine collection, and index five short docs with metadata.
from pathlib import Path
import shutil
import chromadb
DB = Path("chroma_data")
if DB.exists():
shutil.rmtree(DB)
client = chromadb.PersistentClient(path=str(DB))
col = client.get_or_create_collection(
name="pyinns_snippets",
metadata={"hnsw:space": "cosine"},
)
docs = [
"FastAPI JSON Lines streams large responses without loading everything into memory.",
"Polars lazy API defers work until collect so CSV scans stay cheap on big files.",
"ChromaDB stores embeddings and documents for semantic search in RAG pipelines.",
"Instructor wraps LLM calls so responses validate against Pydantic models.",
"DuckDB runs SQL analytics in-process next to Python without a server.",
]
metas = [
{"topic": "fastapi", "kind": "web"},
{"topic": "polars", "kind": "data"},
{"topic": "chromadb", "kind": "rag"},
{"topic": "instructor", "kind": "llm"},
{"topic": "duckdb", "kind": "data"},
]
ids = [f"doc-{i+1}" for i in range(len(docs))]
col.add(ids=ids, documents=docs, metadatas=metas)
print("chromadb", chromadb.__version__, "count", col.count())
get_or_create_collection is idempotent. hnsw:space": "cosine" matches typical text-embedding RAG setups. IDs must be unique strings.
2. Semantic query + metadata filter
Ask in natural language. Chroma embeds the query with the same default model and returns nearest neighbors. Optionally restrict with where on metadata.
query = "How do I do semantic search for RAG with embeddings?"
results = col.query(
query_texts=[query],
n_results=3,
include=["documents", "metadatas", "distances"],
)
for i, (doc_id, doc, meta, dist) in enumerate(
zip(
results["ids"][0],
results["documents"][0],
results["metadatas"][0],
results["distances"][0],
),
1,
):
print(f"#{i} id={doc_id} dist={dist:.4f} topic={meta['topic']}")
print(" ", doc)
filtered = col.query(
query_texts=["analytics on tabular data"],
n_results=2,
where={"kind": "data"},
include=["documents", "metadatas", "distances"],
)
print("filtered kind=data →", [m["topic"] for m in filtered["metadatas"][0]])
3. Real output from our run
On Python 3.13.5 with chromadb 1.5.9, the collection held five documents. The RAG query ranked the ChromaDB snippet first (lowest cosine distance). The kind=data filter returned DuckDB then Polars:
chromadb 1.5.9
client Client
count 5
query: How do I do semantic search for RAG with embeddings?
#1 id=doc-3 dist=0.2465 topic=chromadb kind=rag
ChromaDB stores embeddings and documents for semantic search in RAG pipelines.
#2 id=doc-4 dist=0.8471 topic=instructor kind=llm
#3 id=doc-1 dist=0.9248 topic=fastapi kind=web
filtered where kind=data:
id=doc-5 dist=0.7718 topic=duckdb
id=doc-2 dist=0.9203 topic=polars
collections: ['pyinns_snippets']
4. Practical tips
- Persist early.
EphemeralClient()is fine for tests; usePersistentClientfor anything you reopen tomorrow. - Stable IDs. Hash file paths or content so re-ingest updates instead of duplicating.
- Chunk before add. Split long PDFs/Markdown (Docling helps) — embedding quality collapses on giant blobs.
- Distances, not scores. With cosine space, lower distance is closer. Do not treat the number as a probability.
- Swap embeddings later. Pass a custom
embedding_functionwhen you move from MiniLM to a provider model. - Demo UIs. Wrap
col.queryin Gradio when you want a shareable retriever playground — see our Gradio ML demos tutorial.
Common pitfalls
- Mixing spaces. Do not query a cosine collection with embeddings built for L2 without rebuilding.
- Forgetting
include. Newer clients may omit documents/metadatas unless you list them. - Telemetry surprises. Set
ANONYMIZED_TELEMETRY=Falsein CI if you need a silent offline run. - Huge first download. The default ONNX MiniLM model is ~79 MB — cache it once, then demos are fast.
What to build next
Pipe retrieved chunks into an Instructor-validated answer schema, or optimize prompts with DSPy while Chroma stays the retriever. When you need a browser demo around col.query, Gradio is the shortest path.