Subsystem 01
Document ingestion engine
OCR, layout-aware PDF parsing and semantic chunking
Typical stack
LlamaIndex + Unstructured
Reference architecture
Document Q&A that answers from your own sources with exact citations, respects document permissions, and says so when the sources don't contain the answer.
Design constraints
Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.
Component topology
Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.
Stack topology
Enterprise RAG pipeline: hybrid search and re-ranking
Illustrative reference architecture
Document ingestion engine
OCR, layout-aware PDF parsing and semantic chunking
LlamaIndex + Unstructured
Embedding & vector store
Dense vector index with metadata permission filtering
Qdrant / PostgreSQL pgvector
Sparse keyword engine
BM25 lexical search for exact acronyms, codes and names
Elasticsearch / OpenSearch
Cross-encoder re-ranker
Re-scores the top 50 fused candidates and keeps the best 5
Cohere Rerank / BGE reranker
Subsystem 01
OCR, layout-aware PDF parsing and semantic chunking
Typical stack
LlamaIndex + Unstructured
Subsystem 02
Dense vector index with metadata permission filtering
Typical stack
Qdrant / PostgreSQL pgvector
Subsystem 03
BM25 lexical search for exact acronyms, codes and names
Typical stack
Elasticsearch / OpenSearch
Subsystem 04
Re-scores the top 50 fused candidates and keeps the best 5
Typical stack
Cohere Rerank / BGE reranker
Subsystem 05
Grounded answers with citations, plus PII redaction
Typical stack
Claude API / self-hosted models on vLLM
Data lifecycle
The user submits a query with their enterprise identity and group permissions.
The query becomes a dense embedding and a set of BM25 keyword tokens, in parallel.
Hybrid search queries the vector store and the lexical engine, filtering candidates by the user's permissions.
Reciprocal Rank Fusion (RRF) merges both lists and passes the top 50 candidates to the cross-encoder re-ranker.
The top 5 chunks go into the prompt with strict citation instructions, and the LLM streams its answer.
Reliability and resilience
Failure mode 01
Mitigation
Semantic chunking plus a strict score cutoff on the re-ranker.
Failure mode 02
Mitigation
An automated check (for example in DeepEval) that every number in an answer appears in the retrieved chunks; failing answers are blocked or flagged.
Failure mode 03
Mitigation
Cache embeddings for frequent queries in Redis and run a local embedding model (such as BGE-Large) on dedicated inference GPUs.
Questions
Dense embeddings capture meaning but often miss exact part numbers, acronyms and names; BM25 catches those. Combining the two and re-ranking the result recovers exact matches without losing semantic recall. We measure recall on a labeled set of your own queries rather than quoting a benchmark.
Document ACLs are stored as metadata on each chunk and evaluated at query time against the user's current roles, so a permission change means updating metadata, not re-embedding the document.
Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.