RAGStackGuide
RAG Pipeline Engineering Tool

RAG Chunking Strategy & Vector Index Calculator

Estimate how many vectors a corpus produces at a given chunk size and overlap, what those vectors cost in index RAM, and what one pass of embedding costs at the published rate for the model you pick. Every figure below is arithmetic on the inputs, with the assumptions stated at the bottom of the page.

Corpus and model parameters

50M Tokens
Vector Index RAM
0.99 GB
Vectors plus index overhead
Total Vectors Generated
114,890
Chunk overlap applied
Effective Chunk Stride
435 tokens
Corpus advanced per chunk

RAG index capacity and memory allocation Calculated from the inputs

RAG Pipeline Metric Sizing Value
Bytes per stored vector 6,144 B
Raw embedding vector size 673 MB
Index and metadata overhead (+50%) 337 MB
One-time embedding API cost $1.00
Indicative instance size 2 vCPU / 4 GB RAM (single node)

What these inputs imply

  • Corpus chunked into 114,890 vector embeddings at 512 tokens per chunk.
  • Switching to 8-bit scalar quantisation would cut stored vector bytes roughly fourfold.
  • Estimated vector index RAM: 0.99 GB.

How these numbers are calculated

  • Chunk count is the corpus token total divided by the effective stride, which is chunk size multiplied by one minus the overlap ratio. Overlap re-embeds text, so it raises the vector count rather than leaving it unchanged.
  • Vector bytes are dimensions multiplied by 4 for unquantised float32, or by 1 for 8-bit scalar quantisation.
  • Index RAM applies the 1.5x multiplier Qdrant documents for in-memory collections, where the extra 50% covers index structures, point versions and temporary segments during optimisation. A flat index carries no graph, so no multiplier is applied.
  • Embedding cost follows the published OpenAI list rate for the model selected above: $0.02 per million tokens for text-embedding-3-small and $0.13 for text-embedding-3-large. BGE-M3 and Nomic Embed v1.5 are open-weight and shown at zero because the calculator assumes you run them on your own hardware; a hosted endpoint for either would carry its own rate. This is a one-time indexing cost rather than a monthly charge, and re-indexing after a model change repeats it in full.

These are sizing estimates, not measurements. Payload storage, replication factor and filtering indexes all add to the real footprint, and no retrieval-quality figure is estimated here because recall depends on your corpus and questions, not on these inputs.

Primary sources: Qdrant capacity planning, Qdrant quantization, OpenAI embeddings guide, OpenAI API pricing.

Choosing the inputs

The two inputs that move every other number are chunk size and embedding width. Both are explained in the guides below.