DEV Community

Monalisa Das
Monalisa Das

Posted on

RAG always needs a dedicated vector database — challenged

Hot take: you don't need a vector database for RAG.

Most teams reach for a purpose-built vector store on day one because the onboarding docs tell them to. That's premature infrastructure.

A benchmark on financial documents this year found BM25 keyword search outright beat dense retrieval. Embeddings smear identifiers, version strings, and SKUs into their semantic neighborhood - and lose precision on exact-match queries. That's the failure mode nobody warns you about. Hybrid BM25 + vector with reciprocal rank fusion consistently outperforms either alone. And for most teams under roughly 1M vectors, pgvector + BM25 on the Postgres instance you already run beats standing up a new database you now have to operate.

A dedicated vector DB is a scaling decision you earn - not a starting assumption. What's your retrieval default before you've actually measured quality on your query mix?

https://supabase.com/blog/pgvector-vs-pinecone

Top comments (0)