A developer told me about the first time he built a RAG application. He used Chroma to chunk documents, generate embeddings, and index everything locally. The demo ran smoothly. His manager was impressed. Then the app went live.
Three weeks after launch, query latency started creeping up. One morning, a pod restarted and the vector index was just gone. He spent a weekend tracing the problem and eventually found the answer: Chroma defaults to in-memory storage, persistence requires explicit configuration, and the configuration he had wasn’t quite right.
This wasn’t a bug in Chroma. It was a mismatch between the tool and the use case. Chroma was built for quick validation, and it’s honest about that in its documentation. But when a project moves from demo to production, that mismatch starts to matter.
Vector databases became a hot category almost overnight because large language models made RAG a standard pattern. If your application needs to retrieve relevant context from a document corpus before generating a response, you need somewhere to store and query vector embeddings. That’s the basic job. What’s less obvious is that different tools doing that basic job have very different assumptions about scale, operational complexity, and the kinds of queries you’ll actually run.
This piece walks through Chroma and five serious alternatives: Qdrant, Milvus, Weaviate, pgvector, and Pinecone. The goal isn’t to name a winner. The goal is to help you think through which one fits where you are right now.
The Self-Hosted Open Source Path
For teams that have engineering capacity and want maximum control over their infrastructure, self-hosted open source is the natural direction. Two projects stand out in this space.
Qdrant
Qdrant is written in Rust, which signals its design priorities clearly: low memory overhead, consistent performance, predictable resource consumption. For a database you’re running in production, these properties matter more than they might seem at selection time.
The feature that comes up most in production RAG applications is filtered vector search. In practice, you rarely want to search across all your vectors. You want to find documents similar to a query vector, where those documents also match certain conditions: they belong to a specific user, they’re tagged with a particular category, they were uploaded after a certain date. Chroma handles this, but Qdrant was designed with it as a primary use case. Its payload filtering system is flexible and performant, handling complex boolean conditions without degrading search quality.
Qdrant supports multiple indexing strategies and lets you tune the tradeoff between search speed and recall based on your actual requirements. It has a managed cloud option if you want to avoid running it yourself, but the self-hosted path is well-documented and the Docker deployment is straightforward.
The learning curve is real. Where Chroma feels like a Python library, Qdrant feels like a database. You need to understand its model: collections contain points, each point has a vector and a payload, queries return scored results with payload attached. Once that mental model clicks, the API feels logical. Getting there takes a bit of time.
For projects in the range of hundreds of thousands to tens of millions of vectors, where production stability matters and metadata filtering is part of your query patterns, Qdrant is the most natural step up from Chroma.
Milvus
If Qdrant’s design target is “production-grade vector database,” Milvus aims at something larger: enterprise-scale vector infrastructure. It’s an LF AI & Data Foundation project, developed primarily by Zilliz, and its architecture reflects an assumption that you might be dealing with hundreds of millions or even billions of vectors.
The way Milvus handles this scale is through a distributed architecture with separated compute and storage layers. Each layer scales independently, which means you’re not forced to over-provision the whole system when only one bottleneck needs more capacity. It supports a wide range of indexing algorithms, IVF variants, HNSW, DiskANN, giving you control over the memory/performance tradeoff at a level most teams won’t need but some actually do.
The cost of this capability is operational complexity. A production Milvus deployment involves etcd for coordination, MinIO or similar for object storage, and a message queue component. Getting all of that running reliably requires real infrastructure experience. If your team doesn’t have someone who’s comfortable with distributed systems operations, the gap between “Milvus running” and “Milvus running reliably in production” is wider than the documentation suggests.
There’s a lighter-weight Milvus Lite for local development and small-scale deployment, which helps with the development experience. Zilliz also offers a managed cloud version that removes the operational burden entirely if you want Milvus’s query capabilities without managing its infrastructure.
The honest version of the Milvus pitch is: if you know you’re building something that will eventually deal with hundreds of millions of vectors, or if you’re at an organization that already runs distributed infrastructure and the operational cost is lower for you than it would be for a startup, Milvus is worth serious evaluation. If neither of those is true, it’s probably more than you need.
—
Weaviate: When You Need More Than Vectors
Weaviate approaches the space from a slightly different angle. It positions itself as a vector search engine rather than just a vector database, and the distinction shows up most clearly in two features: hybrid search and the module system.
Hybrid search is the ability to combine vector similarity search with keyword-based full-text search in a single query. This matters more often than it might seem. Pure semantic search is great for conceptual similarity, but it can miss exact terminology matches, product names, or specific identifiers that a user might type verbatim. Pure keyword search misses synonyms and paraphrases. Hybrid approaches blend both signals, which tends to produce better results for real-world search interfaces.
The module system lets you run embedding models directly within Weaviate, so data gets vectorized as it’s inserted rather than requiring an external vectorization step. For teams who want to reduce the number of moving parts in their pipeline, this is appealing.
Weaviate is available as a managed cloud service or as a self-hosted deployment. The community is active and the documentation is thorough. The main friction point is that its API design feels different from other databases in this space: it uses a GraphQL-based query interface, and the schema definition approach is fairly opinionated. Teams migrating from other systems often mention an adjustment period.
For knowledge bases, customer-facing search, or any application where users might type either a conceptual question or a specific term and expect good results either way, Weaviate’s hybrid search capability solves a real problem and isn’t trivial to replicate with other tools.
—
pgvector: The Already-Have-Postgres Option
There’s an often-overlooked scenario worth addressing directly. If your application already runs on PostgreSQL, your first question probably shouldn’t be “which vector database should I add?” It should be “does pgvector cover my needs?”
pgvector is a PostgreSQL extension that adds vector storage and similarity search as native database features. You store vectors in regular Postgres columns, build indexes on them, and query them with SQL. Cosine similarity, Euclidean distance, and inner product are all supported.
The reason this is underrated is everything that comes with being inside PostgreSQL. Your vector data lives in the same database as your application data, so joins work naturally. You can combine vector similarity search with arbitrary SQL conditions in a single query, and those conditions are fully indexed and optimized by the Postgres query planner. Transactions, ACID guarantees, and your existing backup strategy all apply automatically. Your team already knows how to operate Postgres.
The honest limitation is scale. pgvector works well up into the millions of vectors with proper indexing, but if your dataset is tens of millions of rows and growing rapidly, query performance will eventually become a concern. Index build times on very large datasets can be significant. This isn’t a flaw, it’s just the boundary of what Postgres is optimized for.
For applications that are already Postgres-native, have data volumes in a range where Postgres handles it comfortably, and don’t need the specialized indexing algorithms of purpose-built vector databases, pgvector avoids adding an entirely new system to your infrastructure. That simplicity has real value.
—
Pinecone: Paying to Not Think About Infrastructure
The previous options all involve deploying and operating something yourself. Pinecone is a fully managed cloud service, and that’s the whole point.
You create an index, get an API endpoint, and push vectors to it. Pinecone handles everything else: sharding, replication, scaling, failover. Query latency is consistently low. The operational surface from your side is essentially zero. It integrates cleanly with all the major AI and ML frameworks.
For a team whose main priority is shipping AI product features rather than running database infrastructure, this trade makes sense. Engineering time is expensive. The cost of a managed service is predictable and shows up on a bill. The cost of running your own distributed vector database shows up in incidents, on-call rotations, and debugging sessions at inconvenient times.
The considerations that matter when evaluating Pinecone are different from the technical ones. First, cost at scale: managed services price on queries and storage, and those numbers grow with usage in ways that can become significant for high-volume applications. Second, data residency: if your application handles sensitive data or has compliance requirements, storing that data on a third-party service requires careful evaluation. Third, vendor dependency: migrating off Pinecone later involves re-generating all your vectors for a new system, which at large scale is a real project.
None of these are reasons to avoid Pinecone. They’re things to price into the decision upfront rather than discover later.
—
The Comparison at a Glance
The differences between these tools become clearer when you look at a few specific dimensions side by side.
| Tool | Deployment | License | Pricing Model | Best Fit |
|---|---|---|---|---|
| Chroma | Local / self-hosted | Apache 2.0 | Free (open source) | Prototypes, local dev, small data |
| Qdrant | Self-hosted / managed cloud | Apache 2.0 | Free self-hosted; cloud usage-based | Production apps needing filtered search |
| Milvus | Self-hosted / Zilliz cloud | Apache 2.0 | Free self-hosted; cloud usage-based | Billion-scale, enterprise infra |
| Weaviate | Self-hosted / managed cloud | BSD-3 | Free self-hosted; cloud per-node | Hybrid vector + keyword search |
| pgvector | PostgreSQL extension | PostgreSQL License | Free (bundled with Postgres) | Existing Postgres, mid-scale simplicity |
| Pinecone | Fully managed cloud | Commercial (closed) | Free tier + usage-based | Managed infra, fast time-to-production |
A few notes on this table: pricing details change, so treat these as directional rather than precise. All the self-hosted options have associated hosting and operational costs that don’t show up in the license column. The “best fit” descriptions are generalizations; real selection should account for your specific query patterns, not just data volume.
—
The Questions That Actually Drive the Decision
A comparison table helps with initial filtering, but the factors that determine the right choice for a specific team are usually more situational than a table captures.
How much can your team actually operate? This is the most underweighted factor in database selection discussions. Milvus’s distributed architecture is powerful in practice, but power that you can’t reliably maintain is a liability in production. If no one on your team has experience running distributed stateful systems, the gap between deploying Milvus and running Milvus well is real. Qdrant and pgvector are considerably more approachable for smaller teams with generalist infrastructure experience.
What’s your data volume and growth trajectory? There’s a meaningful difference between “we have a few hundred thousand documents now and might grow to a few million” and “we’re ingesting at a rate that will put us in the hundreds of millions within a year.” The first scenario is comfortable for Qdrant or pgvector. The second warrants looking at Milvus or Pinecone from the start.
What does your query pattern actually look like? Pure vector similarity search is the baseline, and every tool in this list handles it. The divergence comes from what you need beyond that baseline. Frequent multi-field metadata filtering points toward Qdrant. Hybrid vector-plus-keyword queries point toward Weaviate. Complex joins with relational application data point toward pgvector. If your query pattern amounts to “give me the N most similar vectors to this embedding,” almost anything will work and you should optimize for operational simplicity instead.
Does data residency matter? Regulatory requirements, contractual commitments, or internal data governance policies can effectively remove fully managed options from consideration. If you need data to stay in specific infrastructure that you control, the choice narrows to self-hosted options.
—
Don’t Underestimate Migration Cost
One more thing worth flagging explicitly: migrating between vector databases is more expensive than it looks on paper.
Vector embeddings are coupled to the model that generated them. If you move from one database to another and also change your embedding model (which happens frequently, because models improve), you need to regenerate every vector in your corpus. At large scale, that’s real compute cost and real calendar time. Even if you keep the same embedding model, migrating the data itself, re-indexing, testing that search quality is equivalent, and updating application code to use the new query API all take engineering time.
This doesn’t mean you should over-engineer your initial selection to avoid all future migration. It means you should be honest with yourself about whether your current choice is a temporary tool or a long-term foundation, and architect accordingly. If you’re using Chroma for a proof-of-concept, build your embedding and retrieval logic behind a thin abstraction layer. When the time comes to move to a production system, you’ll be replacing one implementation rather than untangling assumptions spread across the entire codebase.
If your project already has significant scale, the migration path itself should be a first-class part of the evaluation. Don’t pick a new database and then figure out how to get your data there. Figure out the migration path first, then confirm the destination is worth it.
—
Fitting Tool to Stage
The pattern that emerges from looking at all these options together is less “which database is best” and more “which database fits where the project is right now.”
Chroma is an honest starting point. It’s fast to set up, integrates with everything, and gets out of your way while you’re figuring out whether your RAG approach actually works. The moment production stability or query complexity becomes a requirement, it starts to show its limits, but that’s not a failure mode, it’s just the edge of the design envelope.
Qdrant and pgvector occupy a similar position in the middle: capable enough for real production workloads at significant scale, without requiring a dedicated infrastructure team to keep running. They’re the right move when you’ve validated the approach and need something you can actually rely on.
Milvus handles the high end of the scale spectrum and makes sense when you know from the start that the data volume is going to be large, or when you’re in an organization that already operates distributed infrastructure and the marginal cost of one more system is low.
Pinecone trades cost for time. For teams where the bottleneck is engineering capacity rather than budget, that trade is often sensible. For teams where the data volume is high or the regulatory context is complex, the tradeoffs cut the other way.
The thing to avoid is selection anxiety: spending weeks evaluating tools when Chroma would have gotten you to a working prototype in a day, or conversely, committing to Chroma for a production system because switching feels complicated when it was always going to need to change.
Figure out what stage you’re at, pick the tool that fits it, and leave yourself a clean path to the next one.



