
Vector database pricing pages are built to be read one vendor at a time, and that is the problem. Pinecone bills read units that grow with the size of your namespace. turbopuffer bills petabytes scanned. Chroma bills tebibytes queried. Qdrant and MongoDB bill node hours, Weaviate bills vector dimensions, and Postgres bills an instance. None of those units converts to another without a workload, so this post supplies one and prices it everywhere.
The workload is ten million 1536-dimension embeddings, which is roughly what a mid-sized RAG product holds once real documents are loaded. This post is for the engineer choosing a store before launch, or explaining why the current one costs ten times what the proof of concept did. Every rate was read from the vendor on a stated day, and every monthly total comes from a script printed near the end, so you can re-run it when the prices move.
The short version: query volume decides everything. At one million queries a month, metered turbopuffer costs under 70 USD. At ten million, metered services cost hundreds to thousands, while provisioned clusters from Qdrant, Zilliz and MongoDB sit in the 270 to 420 USD range regardless of traffic. Pinecone’s on-demand tier, which most tutorials start with, is the most expensive option in the table once traffic is real.
Where These Vector Database Pricing Numbers Came From
Every rate below came from the vendor’s own pricing page, pricing documentation, or the calculator embedded in that page. Where a rate only exists inside a calculator, the post says so, because a number with no page text behind it can change without anyone announcing it.
Prices read 2026-09-30
Region: AWS US East (N. Virginia) wherever the vendor lets you choose
AWS RDS rate file published 2026-09-29T23:30:06Z
Pinecone: pinecone.io/pricing and docs.pinecone.io/guides/manage-cost
turbopuffer: turbopuffer.com/pricing and turbopuffer.com/docs/pinning
Chroma: trychroma.com/pricing and docs.trychroma.com/cloud/pricing
Zilliz: zilliz.com/pricing/pricing-guide and docs.zilliz.com
Qdrant: qdrant.tech/pricing and the cloud.qdrant.io calculator API
MongoDB Atlas: mongodb.com/pricing
Weaviate: weaviate.io/pricing, September 2026 price book
Month length: 730 h for AWS and Atlas, 732 h for Qdrant, 720 h for Zilliz
Those hour counts are not a typo. Each vendor converts hourly rates to monthly figures differently, and the totals below use each vendor’s own convention rather than normalizing them away.
The vendor pages are Pinecone pricing, turbopuffer pricing, Chroma Cloud pricing, Zilliz Cloud pricing, Qdrant Cloud pricing, MongoDB Atlas pricing, Weaviate pricing and the RDS for PostgreSQL pricing page. Check any figure against them before signing anything.
Several things are deliberately excluded. Embedding generation, reranking, backups, support plans and committed-use discounts are all out, because each one varies more by contract than by vendor. High availability is out too, with two exceptions forced by the vendors themselves: MongoDB requires at least two search nodes, and Weaviate’s calculator hard-codes three replicas.
The Workload Being Priced
One workload, three traffic levels. Everything that is not query volume stays fixed, so the difference between columns is purely the cost of reads.
- Vectors: 10 million, 1536 dimensions, float32, which is 6,144 bytes each
- Metadata: about 500 bytes per record, so 6,644 bytes per record and 66.44 GB in total
- Layout: one namespace or collection, no multi-tenancy
- Churn: one million upserts a month, each overwriting an existing record
- Queries: 1M, 10M and 100M a month at top_k 10, which averages 0.38, 3.8 and 38 queries per second
- Response size: about 5 KB returned per query, an estimate that assumes metadata comes back and vectors do not
The initial load of ten million records is a one-time cost and is left out of the monthly totals. It ranges from nothing on Zilliz, where bulk import is free, to about 266 USD when upserted into Pinecone in batches.
Four Ways Vendors Bill for Vectors
Before any number makes sense, it helps to see what each vendor actually meters. Vector database pricing splits into four models, and the model predicts the bill far better than the headline rate does.
| Model | Vendors | What grows the bill |
|---|---|---|
| Per query, scaled by namespace size | Pinecone on-demand, turbopuffer, Chroma, Zilliz Serverless | Query count multiplied by data scanned |
| Provisioned nodes | Qdrant, Zilliz Dedicated, MongoDB search nodes, Pinecone Dedicated Read Nodes | RAM needed to hold the index |
| Per vector dimension stored | Weaviate | Vector count multiplied by dimensions |
| General-purpose database | pgvector on RDS | Instance size needed for the index |
The first row is the one that surprises people. In practice, a metered service does not bill you for the ten results you asked for. Instead, it bills you as if each query touched the entire namespace, because in its pricing model it did.
Rate Card: Metered Vector Databases
These are the rates that multiply against query volume. Pinecone publishes ranges that vary by cloud and region, so the table uses the low end, which is also what Pinecone’s own cost estimator defaults to.
| Vendor | Storage | Writes | Reads | Plan minimum |
|---|---|---|---|---|
| Pinecone Standard | 0.33 USD per GB-month | 4.00 to 4.50 USD per million WU | 16 to 18 USD per million RU | 50 USD |
| turbopuffer launch | 0.33 USD per GB-month | 2.00 USD per GB written | 1.00 USD per PB queried | 16 USD |
| Chroma Cloud Team | 0.33 USD per GiB-month | 2.50 USD per GiB written | 0.0075 USD per TiB queried | 250 USD, 100 USD of usage included |
| Zilliz Serverless | 0.30 USD per GB-month | 4 USD per million vCU | 4 USD per million vCU | none |
The read units are where the real definitions live. According to Pinecone’s cost documentation, “a query uses 1 RU for every 1 GB of namespace size, with a minimum of 0.25 RUs per query,” and top_k does not change it. As a result, each query against this 66.52 GB namespace costs 66.52 read units, since Pinecone also counts an 8-byte ID per record.
turbopuffer bills “the actual size of the queried namespace or 1.28 GB, whichever is greater,” with an 80 percent marginal discount between 32 and 128 GB. That discount brings this workload down to 38.89 GB billed per query. Chroma’s documentation and calculator both treat a query as scanning the whole collection, which is 0.0604 TiB here. Meanwhile, Zilliz publishes the answer directly: a read against ten million 1536-dimension vectors costs 75 vCU, so 300 USD per million queries.
Rate Card: Provisioned Clusters
Provisioned options charge for capacity, so the question becomes which size the vendor’s own sizing guidance says will hold ten million 1536-dimension vectors.
| Option | Size chosen | Rate | Monthly |
|---|---|---|---|
| Qdrant Standard, int8 quantized | core1: 4 vCPU, 32 GiB | 0.37344 USD/h | 273.36 |
| Qdrant Standard, float32 in RAM | mx6: 16 vCPU, 128 GiB | 1.49376 USD/h | 1,093.43 |
| Zilliz Dedicated Standard | 3 capacity-optimized CU | 0.175 USD per CU-hour | 378.00 plus storage |
| MongoDB Atlas, binary quantized | 2x S30 plus M10 cluster | 0.24 USD/h per node | 408.80 |
| MongoDB Atlas, scalar quantized | 2x S50 plus M10 cluster | 0.99 USD/h per node | 1,503.80 |
| MongoDB Atlas, float32 | 2x S70 plus M10 cluster | 2.50 USD/h per node | 3,708.40 |
| Pinecone Dedicated Read Nodes | 1x b1, one shard | 336.165 USD per month | 336.17 plus storage and writes |
| Weaviate Flex, cost optimized | HFresh index, calculator | per million dimensions | about 1,968 |
The Qdrant sizes follow the Qdrant capacity planning formulas with 20 percent headroom. Float32 vectors need 70.96 GiB of RAM, which rules out the 64 GiB mx5 and lands on mx6. Quantizing to one byte per dimension, with originals kept on disk, drops the requirement to 19.46 GiB, which fits core1.
MongoDB’s sizing follows its published figure of 6 KB per 1536-dimension vector and its rule that search node RAM should be “at least 10% larger” than the index. That puts float32 on S70 and requires two nodes minimum. The Zilliz row relies on an extrapolation stated plainly: Zilliz publishes capacity-optimized capacity as “8 million 768-dim vectors” per CU, and halving that for 1536 dimensions gives four million per CU, so three CU.
pgvector on RDS Is a Sizing Judgment, Not a Rate
Postgres needs its own section, because the rates are precise and the sizing is not. The pgvector README states that each vector takes 4 * dimensions + 8 bytes, which makes the table 61.52 GB. However, pgvector publishes no formula for HNSW index size, and AWS publishes no guidance on which instance fits ten million vectors.
| Configuration | Instance | Monthly (USD) |
|---|---|---|
| HNSW on binary-quantized expression, re-rank against full vectors | db.r7g.large, 16 GiB, plus 150 GB gp3 | 191.72 |
HNSW on halfvec expression, index held in RAM | db.r7g.2xlarge, 64 GiB, plus 150 GB gp3 | 715.13 |
Both rows are engineering judgments rather than vendor figures, so treat them as a starting estimate. The halfvec index at 2 * dimensions + 8 bytes per vector is about 31 GB before graph overhead, which does not fit comfortably in 32 GiB alongside the buffer cache. Consequently, 64 GiB is the conservative choice. The binary row is cheap because its index is tiny, but re-ranking reads full vectors from disk, so latency depends on the 3,000 IOPS gp3 baseline.
The Monthly Vector Database Pricing Table at Three Traffic Levels
This is the table the rest of the post argues about. All figures are USD per month for the workload above, with rates read on 30 September 2026.
| Option | 1M queries | 10M queries | 100M queries |
|---|---|---|---|
| turbopuffer launch, metered | 67.71 | 419.95 | 3,942.37 |
| pgvector, RDS db.r7g.large, binary | 191.72 | 191.72 | 191.72 |
| Qdrant core1, int8 | 273.36 | 273.36 | 273.36 |
| Zilliz Serverless, eu-central-1 | 330.58 | 3,030.58 | 30,030.58 |
| Zilliz Dedicated, 3 CU | 379.66 | 379.66 | 415.66 |
| MongoDB Atlas, 2x S30, binary | 408.80 | 408.80 | 408.80 |
| Pinecone Dedicated Read Nodes | 411.33 | 411.33 | 1,123.66 |
| Chroma Cloud Team | 639.51 | 4,722.09 | 45,547.92 |
| pgvector, RDS db.r7g.2xlarge, halfvec | 715.13 | 715.13 | 715.13 |
| Qdrant mx6, float32 | 1,093.43 | 1,093.43 | 1,093.43 |
| Pinecone Standard, on-demand | 1,139.49 | 10,718.37 | over rate limit |
| turbopuffer, pinned namespace | 1,266.58 | 1,268.83 | 1,291.33 |
| Weaviate Flex, HFresh | about 1,968 | about 1,968 | about 1,968 |
| MongoDB Atlas, 2x S70, float32 | 3,708.40 | 3,708.40 | 3,708.40 |
Three patterns stand out. First, at low volume the metered services win outright, and turbopuffer wins by a wide margin. Second, by ten million queries every metered service except turbopuffer costs more than every quantized provisioned option. Third, the provisioned rows cluster tightly between 270 and 420 USD once they use quantization, which suggests that is what ten million quantized vectors actually cost to hold in memory.
The Pinecone on-demand cell at 100M is not a price, it is a wall. Pinecone’s rate limits allow 2,000 read units per second per index, and at 66.52 RU per query that caps on-demand at about 30 queries per second. An average of 38 therefore cannot be served, before even considering peaks.
Why Metered Vector Database Pricing Explodes With Traffic
Metered vector databases charge each query for the size of the data it could have searched, not the data it returned. Consequently, a namespace ten times larger makes every query up to ten times more expensive, and query volume multiplies that again. The bill scales with namespace size times traffic, which is the one product that grows fastest after a launch.
That is not a pricing trick. An approximate nearest neighbor search against a larger index genuinely does more work, and serverless vendors pass that through honestly. The trouble is that tutorials and free tiers run against namespaces of a few thousand vectors, where the read unit minimum hides the scaling entirely.
The four metered vendors handle it differently, and the differences are large. Pinecone and Chroma scale linearly with no discount. turbopuffer discounts everything above 32 GB, by 80 percent up to 128 GB and 96 percent beyond, which is why its per-query price at this size is roughly 27 times lower than Pinecone’s. Zilliz Serverless publishes a flat 75 vCU per read at this size, which lands between the two.
Is Pinecone Dedicated Read Nodes Worth It?
Yes, at almost any real traffic level. Dedicated Read Nodes replace read unit charges with a fixed node rate, and one b1 node in aws-us-east-1 costs 336.165 USD a month according to Pinecone’s cost estimator. The break-even point against on-demand reads for this namespace works out as follows:
Pinecone on-demand vs one DRN node: 315,850 queries/month
Above roughly 316,000 queries a month, which is about one query every eight seconds, the dedicated node is cheaper. At ten million queries, on-demand costs 26 times as much as a single dedicated node for the same namespace. Storage, writes and egress bill identically on both, so the whole difference is reads.
The 100M figure uses three replicas, because that is what Pinecone’s estimator suggests for 38 queries per second on this namespace. That replica heuristic comes from the estimator’s code, not from documentation. The Dedicated Read Nodes documentation does confirm that one shard holds up to 250 GB, so this namespace needs only one.
Where turbopuffer Stops Being the Cheapest
turbopuffer is the cheapest option in this post at low traffic, and it stays cheapest longer than intuition suggests. Its fixed monthly cost for this workload is about 28.57 USD, and each million queries adds 39.14 USD. Running those two figures against the flat-rate rows gives three crossover points:
turbopuffer = RDS r7g.large at 4.17M queries/month
turbopuffer = Qdrant core1 at 6.25M queries/month
turbopuffer = Zilliz 3 CU at 8.97M queries/month
Below about four million queries a month, turbopuffer is cheaper than everything here. Between four and nine million, it depends which provisioned option you trust. Above nine million, metered turbopuffer loses to all the quantized clusters.
turbopuffer also offers pinned namespaces, billed at 9.67 USD per GB-month in its calculator with a 128 GB floor. For a 66 GB namespace that floor means paying for nearly twice the data you hold, so pinning only makes sense at this size for latency reasons. The turbopuffer pinning documentation suggests evaluating it when a large namespace sustains more than ten queries per second, and at 38 it costs 1,291.33 USD against 3,942.37 metered.
Quantization Is the Real Price Lever
Look at the provisioned rows again, and the precision column explains more of the spread than the vendor does. Qdrant goes from 1,093 USD at float32 to 273 at int8. MongoDB goes from 3,708 at float32 to 1,504 with scalar quantization and 409 with binary. Postgres goes from 715 to 192.
Therefore, the most important pricing decision for a provisioned vector store is not which vendor, it is how many bytes each dimension occupies in RAM. Qdrant, MongoDB and pgvector all support quantization, and each can keep full-precision originals on disk for re-scoring.
What this post did not measure is recall. Quantization trades accuracy for memory, and the loss depends on your embedding model, your data and your re-scoring setup. Before committing to a quantized configuration, run your own evaluation set through it. A cheaper index that returns the wrong chunks costs more in answer quality than it saves on the bill.
The Numbers This Post Could Not Pin Down
A pricing comparison is only as honest as its gaps, and this one has six worth naming.
- Weaviate: the pricing page lists vector dimensions “from 0.00465 USD per million,” but that is the cheapest optimization profile. The page’s calculator applies higher per-index rates, a three-replica factor and a 1.2 multiplier, so the 1,968 USD figure comes from that formula rather than a quote. Also, the calculator carries only GCP region rates, so an AWS us-east-1 rate was not found.
- Upstash Vector: left out entirely. Ten million vectors at 1536 dimensions is 15.36 billion vector-dimensions, over the 2 billion cap on its pay-as-you-go and fixed plans. The only plan that fits is Pro, which is priced by contacting sales.
- Zilliz capacity: Zilliz publishes per-CU capacity only for 768 dimensions, so the three-CU sizing is an extrapolation. Additionally, Zilliz Serverless is not offered in AWS us-east-1, so its row uses eu-central-1.
- MongoDB region: the search node price table does not state which region it applies to.
- Chroma quota: the default limit is five million records per collection, so this workload needs a quota increase or two collections. Chroma also caps concurrent reads at ten per collection by default, which constrains the 100M column.
- pgvector: there is no published HNSW index size formula, so both Postgres rows are sizing judgments.
A RAG Launch That Got Popular
Consider a small team, three or four engineers, shipping a support assistant over a documentation corpus of several million chunks. During development they choose Pinecone on-demand, because it has no servers to size and a free tier to start on. Traffic in the first weeks is modest, and the bill sits at a few hundred dollars, most of it reads.
Then the assistant gets embedded directly in the product, and query volume climbs by an order of magnitude over a couple of months. The code does not change at all. Nevertheless, the invoice climbs from hundreds to over ten thousand dollars, because every query is still billed against the full namespace and there are now ten times as many of them.
The fix at that point is not necessarily a migration. Moving reads onto one Dedicated Read Node cuts the read line to a flat fee while keeping the same vendor, the same data and the same query API. The mistake was never choosing Pinecone. Instead, it was not recomputing the bill when the traffic profile changed, which is exactly the recomputation the script below makes cheap.
When to Use Each Vector Database Pricing Model
Start on Metered turbopuffer When Traffic Is Unproven
- Your query volume is under about four million a month, where nothing else in this table is cheaper
- Traffic is spiky or seasonal, so provisioned capacity would sit idle most of the month
- You run many namespaces per tenant, each small and rarely queried
- You want to defer sizing decisions until real usage data exists
Move to Provisioned Capacity Once Queries Pass Ten Million
- Traffic is steady and high enough that per-query billing dominates the invoice
- You can quantize vectors and keep originals on disk, which is where provisioned pricing gets cheap
- You need predictable monthly spend for budgeting, since node pricing does not move with traffic
- Qdrant core1 and Zilliz Dedicated both land under 420 USD for this workload at any volume
Keep Vectors in Postgres When You Already Run It
- Your vectors live next to relational data you already join against
- You have a team comfortable tuning Postgres memory and index parameters
- Your query volume is moderate and the index fits in RAM after quantization
- Avoiding a second database matters more than the last few points of recall
Pay for Pinecone When Operations Time Is the Constraint
- Your team has no capacity to run or size infrastructure at all
- You switch to Dedicated Read Nodes as soon as traffic passes a few hundred thousand queries a month
- Enterprise features such as BYOC and private endpoints are requirements rather than preferences
When NOT to Use a Managed Vector Database
- Your corpus is small enough, below a million or so vectors, that an embedded store or a Postgres extension handles it without a new bill
- Your queries are mostly keyword lookups with occasional semantic search, where a search engine with vector support is simpler
- Data residency rules require a region none of these vendors operate in
- Your team already runs a Kubernetes platform and can self-host Qdrant or Milvus on reserved capacity
Common Mistakes with Vector Database Pricing
- Estimating the bill from a proof of concept with a few thousand vectors, where read unit minimums hide how cost scales with namespace size
- Assuming top_k controls query cost, when Pinecone, turbopuffer and Chroma all bill by the data scanned
- Staying on Pinecone on-demand past a few hundred thousand queries a month instead of moving to Dedicated Read Nodes
- Sizing provisioned clusters for float32 without testing whether quantization holds your recall
- Ignoring plan quotas, such as Chroma’s five-million-record collection limit, until the load job fails
- Comparing monthly figures across vendors without checking their hour conventions, which range from 720 to 732
How to Re-Run These Numbers Yourself
Vector database pricing changes often, and a stale price is a wrong answer rather than an old one. The whole table above is one PowerShell script over the rates in this post, and it runs on Windows PowerShell 5.1 or PowerShell 7 with no modules.
# Rates read 2026-09-30. 10M vectors, 1536-dim float32, 500 B metadata,
# 1M overwriting upserts a month, top_k 10, about 5 KB returned per query.
$N = 10e6; $Rec = 1536 * 4 + 500 # 6,644 bytes per record
$GB = $N * $Rec / 1e9; $GiB = $N * $Rec / [Math]::Pow(2, 30)
$Churn = 1e6; $Ret = 5000 / 1e9 # GB returned per query
function PineconeOnDemand($q) { # Standard, low end of the range
$pgb = $GB + $N * 8 / 1e9 # Pinecone counts an 8-byte ID
$pgb * 0.33 + $Churn * 2 * 6.652 / 1e6 * 4 + $q * $pgb / 1e6 * 16 +
[Math]::Max(0.0, $q * $Ret - 100) * 0.10 }
function PineconeDrn($q) { # one b1 node, aws-us-east-1
$pgb = $GB + $N * 8 / 1e9
336.165 + $pgb * 0.33 + $Churn * 2 * 6.652 / 1e6 * 4 +
[Math]::Max(0.0, $q * $Ret - 100) * 0.10 }
function Turbopuffer($q) { # launch, $16 minimum
$queried = 32 + ($GB - 32) * 0.2 # 80% marginal discount 32-128 GB
[Math]::Max(16.0, $GB * 0.33 + $Churn * $Rec / 1e9 * 0.5 * 2 +
$q * $queried / 1e6 * 1 + $q * $Ret * 0.05) }
function TurbopufferPinned($q) { # 128 GB billing floor, 1 replica
128 * 9.67 + $GB * 0.33 + $Churn * $Rec / 1e9 * 0.5 * 2 + $q * $Ret * 0.05 }
function Chroma($q) { # Team: $250 incl. $100 of usage
$use = $GiB * 0.33 + $Churn * $Rec / [Math]::Pow(2, 30) * 2.5 +
$q * $GiB / 1024 * 0.0075 + $q * 5000 / [Math]::Pow(2, 30) * 0.09
250 + [Math]::Max(0.0, $use - 100) }
function ZillizDedicated($q) { # Standard, capacity-optimized, 3 CU
0.175 * 3 * 720 + $GB * 0.025 + [Math]::Max(0.0, $q * $Ret - 100) * 0.09 }
function ZillizServerless($q) { # aws-eu-central-1 only
$q / 1e6 * 300 + $GB * 0.30 + $Churn * ($Rec / 1000 * 0.25 + 1) / 1e6 * 4 }
"record {0} B, {1:N2} GB, {2:N2} GiB" -f $Rec, $GB, $GiB
"{0,-34}{1,12}{2,12}{3,12}" -f '', '1M', '10M', '100M'
foreach ($f in 'PineconeOnDemand','PineconeDrn','Turbopuffer','TurbopufferPinned',
'Chroma','ZillizDedicated','ZillizServerless') {
"{0,-34}{1,12:N2}{2,12:N2}{3,12:N2}" -f $f, (& $f 1e6), (& $f 1e7), (& $f 1e8) }
record 6644 B, 66.44 GB, 61.88 GiB
1M 10M 100M
PineconeOnDemand 1,139.49 10,718.37 106,547.17
PineconeDrn 411.33 411.33 451.33
Turbopuffer 67.71 419.95 3,942.37
TurbopufferPinned 1,266.58 1,268.83 1,291.33
Chroma 639.51 4,722.09 45,547.92
ZillizDedicated 379.66 379.66 415.66
ZillizServerless 330.58 3,030.58 30,030.58
The PineconeDrn line prices a single node, which is why the 100M column in the main table adds two more replicas. Note the 0.0 and 16.0 literals in every [Math]::Max call. With integer literals, PowerShell picks the integer overload and silently rounds the totals, which is exactly the kind of bug that makes a cost model confidently wrong.
To adapt the script, change $N and the 1536 dimension first, since namespace size drives every metered row. Then change the query volumes in the loop to your own traffic from the last thirty days.
Conclusion: Pick the Billing Model Before the Vendor
Vector database pricing at ten million vectors comes down to one question: how many queries a month you will actually serve. Under about four million, metered turbopuffer is the cheapest option by a wide margin. Above ten million, a quantized provisioned cluster from Qdrant, Zilliz or MongoDB holds the bill under 420 USD regardless of traffic, and Pinecone users should move to Dedicated Read Nodes well before that point.
The next step is to pull your real query count and namespace size, then put them into the script above before choosing anything. If you are still deciding on features rather than price, the vector databases comparison covers Pinecone, Weaviate and Chroma architecture, and Pinecone serverless in production RAG walks through the setup this post prices.
For the Postgres route, pgvector for Postgres RAG covers the index choices, and the managed Postgres pricing comparison prices the instance underneath. Dimension count drives every row in this post, so comparing OpenAI, Voyage and Cohere embeddings is worth reading before you pick a model, because a smaller embedding shrinks the bill everywhere at once.
These prices will be re-read and this post updated within three months.