Embedding cost calculator

By Michael Lip. Enter a document set and this page computes the chunk count after overlap, the one-off embedding API bill at published list prices, the raw vector store size in fp32 and in binary quantization, and a monthly hosting estimate. The per-model prices and dimensions below are published values read from the vendors' own pricing and model pages; the chunk and storage arithmetic runs in your browser and nothing is sent to a server.

Estimate an embedding pipeline

Corpus
Model
Storage and queries

Published prices and dimensions

Prices are published list prices per 1M input tokens, checked against the vendors' own pages on 2026-09-28 (Jina publishes no per-token price for jina-embeddings-v4 on its site, so it is not listed). Dimensions are the exact published defaults; Voyage and Cohere accept only specific dimension sizes, while OpenAI and gemini-embedding-001 truncate to smaller prefixes via Matryoshka training. Cohere's embed-v4.0 also supports dimensionality reduction below its listed set. Batch pricing: OpenAI's Batch API halves the list price for asynchronous jobs.

Model$/1M tokensDefault dimsMRL rangeMax input tokensSource
text-embedding-3-small$0.021536256–15368191OpenAI
text-embedding-3-large$0.133072256–30728191OpenAI
voyage-3.5$0.061024256, 512, 1024, 204832000Voyage
voyage-3.5-lite$0.021024256, 512, 1024, 204832000Voyage
voyage-4-large$0.121024256, 512, 1024, 204832000Voyage
embed-v4.0$0.121536256, 512, 1024, 1536128000Cohere
gemini-embedding-001$0.153072128–3072 (recommended 768, 1536, 3072)2048Google

Where the money actually goes

For a typical corpus the one-off embedding bill is small next to what hosting the vectors costs every month. One million 1536-dimension fp32 vectors are 6.1 GB before any index overhead, which puts a Pinecone serverless namespace at about $2/GB-month territory once read units are counted, and makes binary quantization (1 bit per dimension, 32x smaller than fp32 at 1536d) the difference between a hobby project and a line item. The overlap control matters more than people expect: 64 tokens of overlap on 512-token chunks re-embeds 12.5% of the corpus, and that cost is paid on every re-index.

Formula notes

Chunk count is computed with the standard sliding-window formula: chunks = ceil((tokens - overlap) / (chunk - overlap)) per document, where tokens are estimated at 1.33 tokens per word, the ratio OpenAI publishes for English prose. Billing counts actual window lengths, including overlap on every window but never padding beyond the document's estimated tokens. Bytes per vector is dimensions times the precision width (binary is ceil(dimensions/8) packed bytes), multiplied by the index overhead factor for graph-based indexes. The storage figure is the raw vector payload, a floor rather than a provider bill: real invoices also include IDs, metadata, graph memory, retained originals and each provider's minimum resource sizes, and the Qdrant and pgvector figures assume standard tiers and spare instance capacity respectively. Prices are list prices without negotiated discounts, and vector database pricing changes often, so treat the monthly hosting figures as estimates and check the linked price pages.

How we measured this

Every price and dimension in the table above was read from the vendor's own documentation on 2026-09-28: OpenAI's model pages, Voyage AI's published pricing page, Cohere's embed documentation and Google's gemini-embedding-001 page. The chunk and storage arithmetic is plain JavaScript that runs in this browser tab; no number on this page comes from a server call. Michael Lip maintains this page as part of the ml0x calculator set.