Choosing a Vector Database in 2026
pgvector, Pinecone, Weaviate, Qdrant — the trade-offs that actually matter for a production RAG system, with a decision framework.
Choosing a Vector Database in 2026
Every RAG system needs a vector store, and every engineering team goes through the same analysis paralysis. Here is the decision framework I use, updated for the landscape in 2026.
The four questions that matter
Most teams start with benchmark charts. That is the wrong place to start. Begin with these four questions:
- How big is your corpus, really? Under a million chunks, almost anything works. Over 100M, the field narrows fast.
- Do you already run Postgres? If yes,
pgvectoris almost always the right first answer. - What is your update pattern? Batch re-indexing every night is very different from continuous upserts.
- Who will operate it at 3am? A managed service you do not have to wake up for has real, quantifiable value.
pgvector: the boring winner
If you already have Postgres, pgvector removes an entire system from your architecture. HNSW indexes are fast enough for most workloads, and you get transactions, backups, and RLS for free.
CREATE INDEX ON docs USING hnsw (embedding vector_cosine_ops);
The best vector database is the one you already know how to operate. For most teams that is Postgres.
When to reach for a dedicated store
You need a dedicated vector database when:
- Your corpus is large enough that HNSW index build times become a delivery blocker
- You need sub-10ms p99 latency at scale that Postgres cannot sustain
- You want hybrid search features that would require a second system anyway
Here is a quick comparison of the dedicated options I have shipped with:
| Store | Strength | Watch out for | |-------|----------|---------------| | Pinecone | Fully managed, zero ops | Cost at scale, vendor lock-in | | Qdrant | Self-hostable, fast filters | You operate it | | Weaviate | Rich modules, GraphQL API | Complexity surface |
A reference RAG architecture
The recommendation
Start with pgvector. Move to a dedicated store only when you have a measured reason to. I have seen too many teams adopt a shiny vector database early and spend months operating it when a Postgres extension would have served them for years.
Want help with something like this?
I turn articles like this into shipped systems for real teams.


