Skip to content

Irrelevant context, poor recall and wrong answers from RAG

RAG retrieval accuracy and chunking optimization

Improve retrieval with better chunking, hybrid BM25 and vector search, and cross-encoder reranking — measured on an evaluation set built from your own questions.

Symptoms

Signs your platform has this problem

If several of these sound familiar, the plan below is where we would start.

01

Vector search missing key facts

Dense embeddings failing on exact keywords, acronyms, part numbers and table headers.

02

Fragmented context

Fixed-size chunks splitting important sentences across boundaries.

03

Slow search over large stores

Searches over millions of unindexed, full-precision vectors taking seconds.

Remediation plan

How we fix it, step by step

Each phase ends with a measurement, so you can see what changed before the next one starts.

01

Semantic and sentence-window chunking

Embedding small units for precise matching while retrieving the surrounding context.

02

Hybrid search

Combining dense embeddings (for example in Qdrant) with BM25 (for example in OpenSearch) through reciprocal rank fusion.

03

Cross-encoder reranking

Reranking candidates with a cross-encoder or ColBERT so the most relevant chunks reach the model.

04

Quantization and HNSW

HNSW indexes plus 8-bit scalar quantization, which cuts vector memory by about 75%.

Technical checklist

Remediation checklist

What we check before a change goes to production:

  • Replace fixed-length chunking with semantic or sentence-window parsing
  • Deploy hybrid search combining dense vectors with BM25 keyword matching
  • Add cross-encoder reranking before results reach the model
  • Enable HNSW indexing and scalar quantization in the vector database

What we measure

We take a baseline first and report the same measurements after each change, from your own monitoring — evidence, not promised results.

Recall@k
Relevant passages in the top results, on your evaluation set
Answer accuracy
Graded answers on the same evaluation set, before and after
Search latency
Vector and hybrid search time at p95

Related service

AI development

LLM systems built for compliance review: schema-validated extraction, human-in-the-loop workflows, and audit trails — measured in cycle time, not demos.

Explore AI development

Questions

Questions about this remediation

It embeds individual sentences for precise matching, but passes the surrounding sentences to the model so it has enough context to reason.

Scalar quantization stores each dimension as an 8-bit integer instead of a 32-bit float, cutting vector memory by about 75%. Recall usually drops a little, so we measure it on your own queries and rescore top results with full-precision vectors where needed.

Want an engineer to look at this with you?

Send us the symptoms and any metrics you have. We'll reply within one business day, set up a call and agree what to measure before anything changes.