Skip to content

Reference architecture

Enterprise hybrid search engine: dense vectors and BM25

An enterprise search engine that combines the semantic understanding of dense vectors with the exact-match precision of BM25 lexical search.

Design constraints

What the design has to hold to

Targets for the scenario this reference is sized for. A real engagement starts by replacing them with your own numbers.

  • 01Query latency under 50 ms across catalogs of 10M+ documents
  • 02Strong relevance on both natural-language queries and exact SKU or part-number lookups
  • 03Dynamic filtering on facets (price, category, availability, permissions)
  • 04Index updates applied continuously without taking search offline

Component topology

System components and technologies

Subsystems with separate responsibilities, clear contracts between them and storage that scales on its own. The stack named for each is typical, not mandatory.

Stack topology

Enterprise hybrid search engine: dense vectors and BM25

Illustrative reference architecture

  1. 01

    Search query gateway

    Intent classification, spelling correction and parallel dispatch

    FastAPI / Rust API gateway

  2. 02

    Dense vector index

    Semantic and synonym similarity search

    Qdrant / pgvector

  3. 03

    Lexical sparse index

    Exact term matching, SKU lookups and BM25 scoring

    OpenSearch / Typesense

  4. 04

    Reciprocal Rank Fusion (RRF)

    Merges the dense and sparse result lists by rank, with no score normalization

    Rust fusion service

Subsystem 01

Search query gateway

Intent classification, spelling correction and parallel dispatch

Typical stack

FastAPI / Rust API gateway

Subsystem 02

Dense vector index

Semantic and synonym similarity search

Typical stack

Qdrant / pgvector

Subsystem 03

Lexical sparse index

Exact term matching, SKU lookups and BM25 scoring

Typical stack

OpenSearch / Typesense

Subsystem 04

Reciprocal Rank Fusion (RRF)

Merges the dense and sparse result lists by rank, with no score normalization

Typical stack

Rust fusion service

Data lifecycle

End-to-end data flow

  1. A user types a query; the gateway generates an embedding and extracts keyword tokens in parallel.

  2. The gateway queries Qdrant (dense vectors) and OpenSearch (BM25) concurrently.

  3. Each engine returns its top 50 scored candidates with metadata.

  4. RRF computes a unified score for each document: RRF_score = 1 / (60 + rank_vector) + 1 / (60 + rank_BM25).

  5. Merged, deduplicated results go back to the client within the latency budget.

Reliability and resilience

Failure modes and how each is contained

Failure mode 01

Vector search ignoring exact SKU queries

Mitigation

An intent classifier detects alphanumeric codes and routes SKU-pattern queries to BM25 alone.

Failure mode 02

Slow query embedding generation

Mitigation

Quantized embedding models (BGE-small, MiniLM) on dedicated CPU or GPU inference nodes keep embedding time to a few milliseconds.

Failure mode 03

Deletions out of sync between indexes

Mitigation

Stream deletion events through Kafka and remove each record from both the dense and the sparse index.

Questions

What teams ask about this design

RRF combines rankings from several search systems using only each document's rank, so scores on different scales never need normalizing. It gives robust relevance across very different query types.

Shoppers search with phrases ('comfortable summer shoes') and with exact product codes ('NIKE-AIR-90-RED'). Hybrid search handles both in a single query.

Planning a system like this?

Send us your requirements, expected load and budget. We'll reply within one business day with an honest read on the design, and on whether we're the right team to build it.