Free calculator · AI
Plan a retrieval-augmented generation (RAG) project
Answer seven questions about your documents and requirements. The guide lists the pipeline work your project needs, from parsing and chunking to permission filtering, re-ranking and an evaluation plan, and places it in a complexity band from a points table shown in full.
How this is calculated
A decision guide, not a calculator of outcomes. Each answer adds points from a fixed table and switches pipeline considerations on or off; the total places the project in a complexity band. Considerations are listed in pipeline order, from getting text out of documents to evaluating answers.
Step by step
- Score each of the seven answers from the points table: document type, corpus size, update rate, access control, citations, languages and response time.
- Add the points (0 to 17) and place the total in a band: Straightforward 0–5, Moderate 6–10, Complex 11–17.
- For each consideration, decide from the answers whether it applies and whether it is essential, recommended or optional; for example, permission filtering applies only when access is restricted.
- List the considerations that apply in pipeline order: ocr before anything else, layout-aware parsing, chunk on document structure, multilingual embeddings, choose exact or approximate search, hybrid retrieval, permission filtering during search, incremental indexing, re-ranking, fewer, better passages in the prompt, citations that can be checked, a latency budget, an evaluation plan before launch.
Default assumptions
Assumptions marked adjustable can be changed in the calculator; the others are fixed parts of the model.
| Assumption | Default | Sources |
|---|---|---|
| Hardest document type (example answer)adjustable | Tables and complex layouts | None |
| Corpus size (example answer)adjustable | 1,000 to 100,000 | None |
| Update rate (example answer)adjustable | Daily or weekly | None |
| Access control (example answer)adjustable | By team or customer | |
| Citations (example answer)adjustable | Link the source document | None |
| Languages (example answer)adjustable | One language | None |
| Response time (example answer)adjustable | Interactive | |
| Points: hardest document type | Plain text or Markdown 0; Digital PDFs and office files 1; Tables and complex layouts 2; Scans and images 3 | None |
| Points: corpus size | Under 1,000 documents 0; 1,000 to 100,000 1; 100,000 to 1 million 2; Over 1 million 3 | None |
| Points: how often documents change | Rarely 0; Daily or weekly 1; Within minutes 2 | None |
| Points: access control | Everyone sees everything 0; By team or customer 1; Per document or user 3 | None |
| Points: citations in answers | Not needed 0; Link the source document 1; Quote the exact passage or page 2 | None |
| Points: languages | One language 0; Two or three 1; Many, or across languages 2 | None |
| Points: response time | Relaxed 0; Interactive 1; Fast 2 | None |
What this doesn’t model
- The points and bands are a rubric for comparing projects, not a measured scale.
- It predicts no accuracy, hallucination rate, latency or cost; measure those on your own documents and questions.
- It doesn't recommend specific vendors or models.
- It assumes one corpus and one kind of user; mixed workloads may need answering separately.
Sources
- Es, James, Espinosa-Anke and Schockaert (arXiv:2309.15217), Ragas: Automated Evaluation of Retrieval Augmented Generation (26 Sep 2023). Accessed . Metrics for whether retrieval finds relevant, focused passages and whether the model uses them faithfully, without human-annotated ground truth.
- Liu et al. (arXiv:2307.03172), Lost in the Middle: How Language Models Use Long Contexts (6 Jul 2023). Accessed . Performance is often highest when relevant information is at the beginning or end of the context and drops when it is in the middle.
- pgvector (GitHub), pgvector: open-source vector similarity search for Postgres. Accessed . Exact search by default (perfect recall); HNSW and IVFFlat trade some recall for speed. With approximate indexes, WHERE filters apply after the index scan; iterative index scans fetch more. Combine with Postgres full-text search for hybrid search.
- Qdrant, Indexing (Qdrant documentation). Accessed . HNSW vector index extended with edges based on indexed payload values, so filters apply during search; a tenant index builds sub-indexes per tenant.
- Qdrant, Hybrid Queries (Qdrant documentation). Accessed . Combines dense vectors (meaning) with sparse vectors (exact word matching) and fuses result lists with Reciprocal Rank Fusion.
- Sentence Transformers, Retrieve & Re-Rank (Sentence Transformers documentation). Accessed . Retrieve a list of candidates (e.g. 100), then re-rank them with a cross-encoder, which is more accurate because it attends across query and document, but too slow to score a whole corpus.
Last reviewed by the QuantmHill engineering team. Found an error?
Link to or cite this tool
Writing about this topic? Link to the calculator or cite it. Its method, defaults and sources are all on this page, so readers can check the numbers.
Embed this calculator
You can put this calculator on your own site for free. Paste the code below where it should appear. It loads the same calculator in a frame, with a link back to this page for the full method and sources.
The credit line links to this page with the anchor text “QuantmHill”. You may edit it, add rel="nofollow" or remove it — the calculator works the same either way. Add ?theme=light or ?theme=dark to the iframe address to fix its colour scheme; otherwise it follows the visitor's system setting.
Add this once per page, after the iframe, if you want the frame to grow and shrink with the calculator instead of using the fixed height above. It accepts messages from quantmhill.com only and resizes only the frame that sent them.
Frequently asked questions
Which parts of a retrieval-augmented generation pipeline your project needs and how strongly (essential, recommended or optional), listed in pipeline order, and a complexity band from a points table you can read in full. It doesn't predict accuracy, hallucination rates or infrastructure cost; those depend on your documents and questions and have to be measured.
Each of the seven answers adds points from the table on this page, out of 17 in total. The bands are: Straightforward 0–5, Moderate 6–10, Complex 11–17. The points are our own rubric for comparing projects, not a measured scale.
Because permissions have to be copied from source systems, kept in sync and applied during the search itself. The pgvector documentation notes that with approximate indexes a WHERE filter is applied after the index scan, which can return too few results, and Qdrant's documentation describes building payload indexes so filters apply during search.
Hybrid search combines keyword matching with vector search, so exact terms such as product codes are still found; Qdrant and pgvector both document ways to fuse the two. Re-ranking runs a slower, more accurate cross-encoder over a shortlist, as the Sentence Transformers documentation describes. The guide makes them essential for large corpora or passage-level citations.
Longer prompts cost more and are slower, and 'Lost in the Middle' (Liu et al., 2023) found that models often use information at the start and end of a long context better than information in the middle. Sending fewer, better-ranked passages usually works better.
Collect real questions with known good sources, then measure two things separately: whether retrieval finds the right passages, and whether answers stay faithful to them. The Ragas paper (Es et al., 2023) proposes metrics for both that don't need human-written reference answers. Re-run the set after every change.
Want an engineer to check your numbers?
Send us your inputs and the decision you're weighing. We'll reply within one business day with an honest read on whether we can help.