Skip to content

Reference architecture

Enterprise RAG pipeline: hybrid search and re-ranking

Document Q&A that answers from your own sources with exact citations, respects document permissions, and says so when the sources don't contain the answer.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01First token streamed within 2 seconds, retrieval included
  • 02No retrieval of documents the user isn't permitted to read
  • 03Exact source citation (document name, page number, bounding box) for every claim
  • 04High recall over complex multi-page tables and financial filings

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

Enterprise RAG pipeline: hybrid search and re-ranking

Illustrative reference architecture

  1. 01

    Document ingestion engine

    OCR, layout-aware PDF parsing and semantic chunking

    LlamaIndex + Unstructured

  2. 02

    Embedding & vector store

    Dense vector index with metadata permission filtering

    Qdrant / PostgreSQL pgvector

  3. 03

    Sparse keyword engine

    BM25 lexical search for exact acronyms, codes and names

    Elasticsearch / OpenSearch

  4. 04

    Cross-encoder re-ranker

    Re-scores the top 50 fused candidates and keeps the best 5

    Cohere Rerank / BGE reranker

Subsystem 01

Document ingestion engine

OCR, layout-aware PDF parsing and semantic chunking

Typical stack

LlamaIndex + Unstructured

Subsystem 02

Embedding & vector store

Dense vector index with metadata permission filtering

Typical stack

Qdrant / PostgreSQL pgvector

Subsystem 03

Sparse keyword engine

BM25 lexical search for exact acronyms, codes and names

Typical stack

Elasticsearch / OpenSearch

Subsystem 04

Cross-encoder re-ranker

Re-scores the top 50 fused candidates and keeps the best 5

Typical stack

Cohere Rerank / BGE reranker

Subsystem 05

LLM generation & guardrails

Grounded answers with citations, plus PII redaction

Typical stack

Claude API / self-hosted models on vLLM

Data lifecycle

End-to-end data flow

  1. The user submits a query with their enterprise identity and group permissions.

  2. The query becomes a dense embedding and a set of BM25 keyword tokens, in parallel.

  3. Hybrid search queries the vector store and the lexical engine, filtering candidates by the user's permissions.

  4. Reciprocal Rank Fusion (RRF) merges both lists and passes the top 50 candidates to the cross-encoder re-ranker.

  5. The top 5 chunks go into the prompt with strict citation instructions, and the LLM streams its answer.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Context window filled with irrelevant chunks

Mitigation

Semantic chunking plus a strict score cutoff on the re-ranker.

Failure mode 02

Unsupported statistics in answers

Mitigation

An automated check (for example in DeepEval) that every number in an answer appears in the retrieved chunks; failing answers are blocked or flagged.

Failure mode 03

Slow embedding API calls

Mitigation

Cache embeddings for frequent queries in Redis and run a local embedding model (such as BGE-Large) on dedicated inference GPUs.

Questions

What teams ask about this design

Dense embeddings capture meaning but often miss exact part numbers, acronyms and names; BM25 catches those. Combining the two and re-ranking the result recovers exact matches without losing semantic recall. We measure recall on a labeled set of your own queries rather than quoting a benchmark.

Document ACLs are stored as metadata on each chunk and evaluated at query time against the user's current roles, so a permission change means updating metadata, not re-embedding the document.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.