← Back to Tutorials
TutorialFor: AI Engineers, ML Engineers, Platform Engineers, AI Systems Architects

How to Build a Hybrid Search Service with BM25 and Docker

Fuse BM25 with dense retrieval, tune the fusion weight on real relevance judgments, and ship the whole thing as a container that shows you every stage's score.

Updated
#tutorial#intermediate#hybrid-search#bm25#dense-retrieval#reciprocal-rank-fusion#cross-encoder-reranking#fastapi#docker

What you'll build: a hybrid search API that shows its working

By the end you will have a hybrid search service, BM25 plus dense retrieval, fused with a weight you chose from measurement rather than from a blog post, running in a container that exposes every stage's score. Intermediate level, about 90 minutes. When a result looks wrong, you can see which arm put it there.

Most of those 90 minutes are waiting rather than typing: roughly 25 minutes to build the index in step 3, 7 minutes to measure reranking in step 9, and 10 minutes for the first Docker build. Everything runs on CPU. There are no API keys and no accounts anywhere.

Here is the finished service answering a query, running inside its container. The cut keeps the long titles from wrapping on this page:

bash
curl -s "localhost:8000/search?q=does+vitamin+d+reduce+cancer+risk&top_k=2" \  | python -m json.tool | cut -c1-88
text
{    "query": "does vitamin d reduce cancer risk",    "fusion": "weighted",    "alpha": 0.7,    "results": [        {            "doc_id": "MED-4570",            "title": "Vitamin D supplement doses and serum 25-hydroxyvitamin D in the ra            "bm25": 4.3297,            "dense": 0.8108,            "fused": 0.9335,            "rerank": null        },        {            "doc_id": "MED-2762",            "title": "Vitamin and mineral supplements in the primary prevention of cardi            "bm25": 3.8309,            "dense": 0.8282,            "fused": 0.9327,            "rerank": null        }    ]}

Those four numbers per hit are the point of the exercise. The second document wins on the dense arm and loses on BM25, and the fused score says which of the two mattered more.

And here is the measurement that sets alpha to 0.7, produced by the evaluation script you will write in step 7. nDCG@10 scores the top ten results on how many relevant documents they contain and how high up those documents sit, where 1.0 is a perfect ranking:

text
arm                 nDCG@10  Recall@10bm25                 0.3225     0.1471dense                0.3387     0.1580rrf k=60             0.3530     0.1660weighted a=0.7       0.3711     0.1737

Verified against Python 3.13.9, bm25s==0.3.11, PyStemmer==3.1.0, sentence-transformers==6.1.0, numpy==2.5.3, ir_datasets==0.6.3, fastapi==0.141.1, uvicorn==0.53.0, with torch==2.14.0+cpu, transformers==5.17.0 and starlette==1.6.0 as resolved by those pins, on 2026-09-20. Every command and every block of output on this page came from a real run on a laptop CPU, including the full 323-query evaluation and the Docker build.

Prerequisites: Python 3.13, about 3.5 GB of disk, and Docker

You need Python 3.12 or newer (3.13.9 here) and Docker for the last step. Budget roughly 3.5 GB of free disk: about 1 GB for the virtual environment, 215 MB for the two models, and 2.4 GB for the finished image, plus a corpus cache in your home directory. No GPU. The model downloads from Hugging Face are anonymous.

You should already know what an embedding is and roughly what BM25 does. You do not need to have used bm25s, sentence-transformers or FastAPI before.

Create the project and a virtual environment:

bash
mkdir hybrid-search && cd hybrid-searchpython -m venv .venv

Activate it. On Windows:

bash
.venv\Scripts\activate

On macOS or Linux:

bash
source .venv/bin/activate

Write requirements.txt:

text
bm25s==0.3.11PyStemmer==3.1.0sentence-transformers==6.1.0numpy==2.5.3ir_datasets==0.6.3fastapi==0.141.1uvicorn==0.53.0

Install them:

bash
pip install -r requirements.txt

sentence-transformers pulls in torch and transformers, which is most of the download.

Verify the install before step 1. This is the cheapest place to find out something is wrong:

bash
python -c "import bm25s,Stemmer,numpy,sentence_transformers as s;print(s.__version__)"

Expected output:

text
6.1.0

If that prints 6.1.0, every import in this tutorial will resolve.

Run every command from this project root, with the virtual environment active: each script resolves index/ relative to the working directory. The curl lines later, and the backslash-continued commands, assume a POSIX shell. On Windows use Git Bash or WSL, because PowerShell's curl is an alias for Invoke-WebRequest and rejects -s.

How BM25, dense retrieval, fusion and reranking fit together

A hybrid search service is four stages, and the interesting problems all live between them.

BM25 scores a document by how many rare query terms it contains, adjusted for document length. Its scores are unbounded, and their size depends on the query as much as on the document: in step 4 you will see one query top out near 4.9, and a rarer query would score higher. Two documents from different queries cannot be compared.

Dense retrieval encodes the query and each document into a vector and ranks by cosine similarity. With normalised vectors those scores sit between -1 and 1, and in step 4 you will see them land between about 0.4 and 0.8. They mean something different from BM25 scores: not "shares rare words with the query" but "sits near the query in the embedding space".

Fusion is the step that exists because those two number lines cannot be added. There are two standard ways out of it:

  • Reciprocal rank fusion throws the scores away and keeps only the positions. A document at rank r contributes 1/(k + r) from each list. It needs no tuning and no normalisation. Every search engine offers it as a default for that reason.
  • Weighted fusion rescales each arm to the range 0 to 1, then takes alpha of the dense score plus 1 - alpha of the BM25 score. It keeps the score magnitudes, and it gives you one knob to tune. That knob is the thing most tutorials set by feel. You will set it by measurement.

If what you want is to compare the retrievers rather than deploy them, the companion five-arm bake-off harness measures the same trade-off on your own corpus, without any of the serving code below.

Reranking is the optional fourth stage. A cross-encoder reads the query and one candidate document together, in a single forward pass, and scores the pair directly. That is more accurate in principle. It also costs one model call per candidate instead of one per query, and whether the trade pays is a question about your corpus. Step 9 measures it rather than assuming it.

The corpus is NFCorpus, part of the BEIR benchmark suite: 3,633 medical abstracts, 323 test queries written in ordinary English, and human relevance judgments that grade each document 1 or 2 rather than just relevant or not. It is small enough to index on a laptop. Its queries mix precise medical vocabulary with plain questions, and that mixture is exactly what hybrid search is supposed to handle.

Step 1: Load the NFCorpus dataset and inspect it

Goal. Download NFCorpus and see what the documents, queries and judgments actually contain.

Why this step. Every later number depends on the shape of this data, and two of its properties decide design choices you will make in a moment: the documents are long enough to hit the embedder's token limit, and the judgments are graded rather than binary. Looking now costs a minute. Finding out in step 7 costs a rewrite.

Code. Create explore.py:

python
import ir_datasetsds = ir_datasets.load("beir/nfcorpus/test")docs = list(ds.docs_iter())queries = {q.query_id: q.text for q in ds.queries_iter()}judged = {}for qrel in ds.qrels_iter():    if qrel.relevance > 0:        judged.setdefault(qrel.query_id, {})[qrel.doc_id] = qrel.relevanceprint(f"documents {len(docs)}  queries {len(queries)}")first = docs[0]print(f"doc {first.doc_id} title: {first.title[:60]}")print(f"doc words: {len((first.title + ' ' + first.text).split())}")some = sorted(judged)[0]print(f"query {some}: {queries[some]}")print(f"its judgments: {judged[some]}")

The qrels in qrels_iter is the standard information retrieval term for query relevance judgments: which documents a human marked relevant for which query.

Run it.

bash
python explore.py

The first run downloads and unpacks the dataset, which took about a minute here and prints [INFO] lines to stderr while it works.

Expected output.

text
documents 3633  queries 323doc MED-10 title: Statin Use and Breast Cancer Survival: A Nationwide Cohort Sdoc words: 263query PLAIN-1008: deafnessits judgments: {'MED-4532': 1, 'MED-4533': 1, 'MED-4534': 1, 'MED-4535': 1, 'MED-4536': 1}

What just happened. You have 3,633 documents of a few hundred words each, 323 queries, and judgments that carry a grade rather than a yes or no. This particular query's five judged documents all scored 1; elsewhere in the file documents score 2, and step 7's metric treats that difference as real.

Both properties named above show up in this output. The documents are long enough to matter against the embedder's 512-token limit. And the queries are not uniform: deafness is one word, while others are full questions. BM25 handles the first kind well, dense retrieval the second.

Step 2: Build the BM25 index

Goal. Tokenise the corpus and build a BM25 index you can save to disk.

Why this step. BM25 costs nothing to run and is hard to beat on precise vocabulary, so it is the baseline everything else has to justify itself against. Building it first also settles the tokenisation. That is where the quiet bugs live, and there is more on what matters when you run BM25 in production than fits in this step. Stem the corpus but not the query, and "supplements" in the query never matches "supplement" in the document; the arm then underperforms without ever failing.

Code. Create build_index.py:

python
"""Download NFCorpus and build both indexes into index/."""import jsonimport pathlibimport sysimport timeimport bm25simport ir_datasetsimport numpy as npimport Stemmerfrom sentence_transformers import SentenceTransformerINDEX = pathlib.Path("index")EMBEDDER = "BAAI/bge-small-en-v1.5"def load_corpus():    ds = ir_datasets.load("beir/nfcorpus/test")    docs = [        {"id": d.doc_id, "title": d.title, "text": (d.title + " " + d.text).strip()}        for d in ds.docs_iter()    ]    queries = {q.query_id: q.text for q in ds.queries_iter()}    qrels = {}    for qrel in ds.qrels_iter():        if qrel.relevance > 0:            qrels.setdefault(qrel.query_id, {})[qrel.doc_id] = qrel.relevance    return docs, queries, qrelsdef main():    INDEX.mkdir(exist_ok=True)    docs, queries, qrels = load_corpus()    print(f"documents {len(docs)}  queries {len(queries)}  judged queries {len(qrels)}")    stemmer = Stemmer.Stemmer("english")    tokens = bm25s.tokenize(        [d["text"] for d in docs], stopwords="en", stemmer=stemmer, show_progress=False    )    retriever = bm25s.BM25()    retriever.index(tokens, show_progress=False)    retriever.save(str(INDEX / "bm25"), show_progress=False)    print(f"bm25 vocabulary {len(retriever.vocab_dict)} terms")if __name__ == "__main__":    main()

The docstring says both indexes and several imports go unused, because step 3 extends this same file with the dense half.

Run it.

bash
python build_index.py

Expected output.

text
documents 3633  queries 323  judged queries 323bm25 vocabulary 19117 terms

What just happened. There is now an index/bm25/ directory holding the sparse matrix, the parameters and the vocabulary. The vocabulary is the part that matters: bm25s saves vocab.index.json alongside the matrix, and that is what lets a query at serving time map its tokens onto the same term IDs the corpus was indexed with. The stemmer and the stopword list are not saved. The search module you write in step 4 has to recreate both with exactly the same settings.

Step 3: Build the dense vector index with sentence-transformers

Goal. Encode every document into a normalised vector and save the matrix next to the BM25 index.

Why this step. This is the arm that finds documents sharing no words with the query. It is also the expensive one: encoding is a one-time cost you pay at build time so that a query only has to encode itself.

Normalise the vectors. A dot product is then the cosine similarity, and the search code needs no division. Check the token limit. bge-small-en-v1.5 reads 512 tokens, and the print statement below shows you that number rather than leaving it implicit. Documents longer than that are cut off silently, with no warning. A 263-word abstract fits. A 40-page contract would lose almost all of itself, and that is the point at which you chunk documents instead of encoding them whole.

Code. No new imports are needed: build_index.py already imports everything this version of main() uses. Extend main() so that the whole function now reads:

python
def main():    INDEX.mkdir(exist_ok=True)    docs, queries, qrels = load_corpus()    print(f"documents {len(docs)}  queries {len(queries)}  judged queries {len(qrels)}")    stemmer = Stemmer.Stemmer("english")    tokens = bm25s.tokenize(        [d["text"] for d in docs], stopwords="en", stemmer=stemmer, show_progress=False    )    retriever = bm25s.BM25()    retriever.index(tokens, show_progress=False)    retriever.save(str(INDEX / "bm25"), show_progress=False)    print(f"bm25 vocabulary {len(retriever.vocab_dict)} terms")    model = SentenceTransformer(EMBEDDER)    print(f"embedder {EMBEDDER} max_seq_length {model.max_seq_length}")    started = time.perf_counter()    embeddings = model.encode_document(        [d["text"] for d in docs],        normalize_embeddings=True,        batch_size=32,        show_progress_bar=False,    )    elapsed = time.perf_counter() - started    np.save(INDEX / "embeddings.npy", embeddings.astype(np.float32))    print(f"embeddings {embeddings.shape}")    # Elapsed times go to stderr: they differ on every machine, so keeping them    # out of stdout means the checkpoint above is the same for every reader.    print(f"encoded in {elapsed:.0f}s", file=sys.stderr)    (INDEX / "docs.jsonl").write_text(        "\n".join(json.dumps(d) for d in docs), encoding="utf-8"    )    (INDEX / "eval.json").write_text(        json.dumps({"queries": queries, "qrels": qrels}), encoding="utf-8"    )    print("wrote", ", ".join(sorted(p.name for p in INDEX.iterdir())))

encode_document is the retrieval-specific entry point in sentence-transformers 6.x. It behaves like encode but applies the model's document prompt where the model defines one, and its counterpart encode_query does the same for queries.

Run it.

bash
python build_index.py

This is the slow step. It downloads the model on the first run, roughly 130 MB, then encodes 3,633 documents. Three runs here took 1,095, 1,129 and 1,849 seconds, so budget 20 to 30 minutes and leave it alone.

Expected output.

text
documents 3633  queries 323  judged queries 323bm25 vocabulary 19117 termsembedder BAAI/bge-small-en-v1.5 max_seq_length 512embeddings (3633, 384)wrote bm25, docs.jsonl, embeddings.npy, eval.json

The elapsed time goes to stderr instead, alongside the [INFO] lines from ir_datasets, so your terminal also shows something like encoded in 1095s. Every script in the next steps follows the same split. Timings belong on stderr because they differ on every machine. Stdout then stays a checkpoint you can compare against this page character for character.

What just happened. index/ now holds everything the service needs and weighs about 15 MB: the BM25 matrix, a 3,633 by 384 float32 embedding matrix, the documents themselves, and the queries and judgments for evaluation. Nothing downstream ever touches ir_datasets again. That is why the container in step 12 can leave it out.

Step 4: Score a query with BM25 and dense retrieval side by side

Goal. Write the search module that turns a query string into two full score vectors, one per arm.

Why this step. Everything after this is arithmetic on these two vectors. Returning the score for every document, rather than a top-k list, is a deliberate choice: fusing top-k lists means deciding what to do about documents that one arm returned and the other did not, and at 3,633 documents you can simply avoid that question. At ten million documents you cannot, and then you fuse truncated lists and treat a missing document as rank infinity.

Code. Create search.py:

python
"""Hybrid search over the prebuilt indexes: BM25, dense, fusion, optional rerank."""import jsonimport pathlibimport bm25simport numpy as npimport Stemmerfrom sentence_transformers import CrossEncoder, SentenceTransformerINDEX = pathlib.Path("index")EMBEDDER = "BAAI/bge-small-en-v1.5"RERANKER = "cross-encoder/ms-marco-MiniLM-L6-v2"class HybridSearch:    def __init__(self, index_dir=INDEX, load_reranker=False):        index_dir = pathlib.Path(index_dir)        lines = (index_dir / "docs.jsonl").read_text(encoding="utf-8")        self.docs = [json.loads(line) for line in lines.splitlines()]        self.bm25 = bm25s.BM25.load(str(index_dir / "bm25"), show_progress=False)        self.embeddings = np.load(index_dir / "embeddings.npy")        self.stemmer = Stemmer.Stemmer("english")        self.embedder = SentenceTransformer(EMBEDDER)        self.reranker = None    def bm25_scores(self, query):        tokens = bm25s.tokenize(            query,            stopwords="en",            stemmer=self.stemmer,            return_ids=False,            show_progress=False,        )        return self.bm25.get_scores(tokens[0])    def dense_scores(self, query):        vector = self.embedder.encode_query(query, normalize_embeddings=True)        return self.embeddings @ vectorif __name__ == "__main__":    engine = HybridSearch()    bm25 = engine.bm25_scores("does vitamin d reduce cancer risk")    dense = engine.dense_scores("does vitamin d reduce cancer risk")    print(f"bm25  min {bm25.min():.4f}  max {bm25.max():.4f}")    print(f"dense min {dense.min():.4f}  max {dense.max():.4f}")    print("bm25 top:  ", engine.docs[int(bm25.argmax())]["id"])    print("dense top: ", engine.docs[int(dense.argmax())]["id"])

Three things in that constructor do nothing yet. load_reranker, the CrossEncoder import and the self.reranker = None placeholder are all wired up in step 9. Typing them now means the constructor signature never has to change.

Note return_ids=False in bm25_scores. It makes bm25s.tokenize hand back token strings rather than integer IDs, which is what get_scores maps through the saved vocabulary. The stopword list and the stemmer match the ones used at index time, exactly as warned in step 2.

Run it.

bash
python search.py

Expected output.

text
bm25  min 0.0000  max 4.9395dense min 0.4150  max 0.8282bm25 top:   MED-5186dense top:  MED-2762

What just happened. You have both arms working, and the output shows the problem fusion exists to solve. The two arms disagree about the best document, and their score ranges do not overlap in any meaningful way: adding 4.9395 to 0.8282 would be adding a word-overlap statistic to a cosine similarity.

Step 5: Fuse BM25 and dense scores with RRF and weighted fusion

Goal. Add reciprocal rank fusion and weighted min-max fusion to search.py.

Why this step. These are the two fusion functions you will measure against each other in step 7. Writing both now, as plain functions over score vectors, keeps the choice a parameter instead of a rewrite.

The minmax helper has one guard worth explaining. A query with no term in the vocabulary at all produces the same score everywhere, and the usual formula then divides by zero. Mapping that case to all zeros is the honest answer: a constant list carries no ranking information, so it should contribute none.

Code. Add these four functions to search.py, above the HybridSearch class:

python
def minmax(scores):    """Rescale one query's scores to [0, 1]. Constant input maps to all zeros."""    low, high = float(scores.min()), float(scores.max())    if high <= low:        return np.zeros_like(scores)    return (scores - low) / (high - low)def ranks_from(scores):    """1-based rank of every document: the top-scoring document gets rank 1."""    return np.argsort(np.argsort(-scores)) + 1def rrf(bm25_scores, dense_scores, k=60):    """Reciprocal rank fusion (Cormack et al., SIGIR 2009)."""    return 1.0 / (k + ranks_from(bm25_scores)) + 1.0 / (k + ranks_from(dense_scores))def weighted(bm25_scores, dense_scores, alpha):    """Convex combination of min-max normalised scores. alpha weights the dense arm."""    return alpha * minmax(dense_scores) + (1.0 - alpha) * minmax(bm25_scores)

Then replace the __main__ block at the bottom of search.py with:

python
if __name__ == "__main__":    engine = HybridSearch()    query = "does vitamin d reduce cancer risk"    bm25 = engine.bm25_scores(query)    dense = engine.dense_scores(query)    for name, fused in (("rrf", rrf(bm25, dense)),                        ("weighted a=0.7", weighted(bm25, dense, 0.7))):        top = int(fused.argmax())        print(f"{name:<15} top {engine.docs[top]['id']}  fused {fused[top]:.4f}")

k=60 comes from the paper that introduced reciprocal rank fusion, where it was "fixed during a pilot investigation and not altered during subsequent validation". The same paper reports that the choice was near-optimal but not critical. Do not treat 60 as a law: Qdrant's implementation defaults to k=2 over zero-based ranks, while Elasticsearch and OpenSearch default to 60 over one-based ranks. If you move this code to an engine, check which convention it uses before comparing numbers, and compare the engines themselves before picking one.

Run it.

bash
python search.py

Expected output.

text
rrf             top MED-4570  fused 0.0312weighted a=0.7  top MED-4570  fused 0.9335

What just happened. Both fusion methods now agree on a document that neither arm ranked first on its own. That is the behaviour hybrid search is sold on. The two fused scores are still not comparable to each other: RRF scores are sums of two small reciprocals and land near 1/60, while weighted scores live between 0 and 1. Whenever you show a fused score to a user, label which method produced it.

Step 6: Return ranked results with every stage's score

Goal. Add a search method that returns the top-k documents with their BM25, dense and fused scores.

Why this step. The API calls this method, and the per-stage scores are the reason the service is worth building rather than taking off a shelf. When a result looks wrong, you want to see whether the dense arm dragged it in or BM25 did.

Code. Add this method to HybridSearch in search.py, after dense_scores:

python
    def search(self, query, alpha=0.7, fusion="weighted", top_k=5, rerank=0):        bm25 = self.bm25_scores(query)        dense = self.dense_scores(query)        fused = rrf(bm25, dense) if fusion == "rrf" else weighted(bm25, dense, alpha)        depth = max(top_k, rerank)        candidates = list(np.argsort(-fused)[:depth])        rerank_scores = {}        results = []        for position in candidates[:top_k]:            results.append(                {                    "doc_id": self.docs[position]["id"],                    "title": self.docs[position]["title"],                    "bm25": round(float(bm25[position]), 4),                    "dense": round(float(dense[position]), 4),                    "fused": round(float(fused[position]), 4),                    "rerank": round(rerank_scores[position], 4)                    if position in rerank_scores                    else None,                }            )        return results

rerank_scores stays empty until step 9 fills it. The rerank key is in the response from the start, so the shape of a result never changes with configuration. The front end then has one case to handle instead of two.

Replace the __main__ block again, for the last time:

python
if __name__ == "__main__":    engine = HybridSearch()    for hit in engine.search("does vitamin d reduce cancer risk", top_k=3):        print(            f"{hit['doc_id']:<9} bm25 {hit['bm25']:>7.4f}  dense {hit['dense']:>6.4f}"            f"  fused {hit['fused']:>6.4f}  {hit['title'][:34]}"        )

Run it.

bash
python search.py

Expected output.

text
MED-4570  bm25  4.3297  dense 0.8108  fused 0.9335  Vitamin D supplement doses and serMED-2762  bm25  3.8309  dense 0.8282  fused 0.9327  Vitamin and mineral supplements inMED-4645  bm25  4.2738  dense 0.8069  fused 0.9234  Vitamin and Mineral Use and Risk o

What just happened. The engine is complete for retrieval. Look at the first two rows: the second document scores higher on the dense arm and lower on BM25, and it lands second because at alpha 0.7 the BM25 gap outweighs the dense gain. That is the kind of decision that is invisible in a normal search API, and it is about to become the thing you tune.

Step 7: Measure BM25, dense, RRF and weighted fusion with nDCG and Recall

Goal. Score BM25, dense, RRF and eleven values of alpha against the corpus's own relevance judgments.

Why this step. Up to here you have been reading result lists and nodding. That is how people end up shipping alpha=0.5 because it sounds balanced. The judgments that ship with NFCorpus turn the question into arithmetic.

nDCG@10 uses the graded labels. A document labelled 2 contributes more gain than one labelled 1, and a correct answer at rank 1 counts for more than the same answer at rank 10. Recall@10 ignores order and asks only how many of the known relevant documents made the cut. The two can disagree. When they do, work out which of them your users actually feel.

Code. Create evaluate.py:

python
"""Score every retrieval arm on NFCorpus, then tune alpha with cross-validation."""import jsonimport mathimport pathlibimport sysimport timeimport numpy as npfrom search import HybridSearch, rrf, weightedINDEX = pathlib.Path("index")ALPHAS = [round(a, 1) for a in np.arange(0.0, 1.01, 0.1)]def ndcg_at_k(ranked_ids, judged, k=10):    """Graded nDCG: NFCorpus labels documents 1 or 2, not just relevant or not."""    gain = sum(        (2 ** judged.get(doc_id, 0) - 1) / math.log2(rank + 2)        for rank, doc_id in enumerate(ranked_ids[:k])    )    ideal = sum(        (2**label - 1) / math.log2(rank + 2)        for rank, label in enumerate(sorted(judged.values(), reverse=True)[:k])    )    return gain / ideal if ideal else 0.0def recall_at_k(ranked_ids, judged, k=10):    return len(set(ranked_ids[:k]) & set(judged)) / len(judged)def score_matrix(engine, scores, query_ids, qrels, k=10):    ndcgs, recalls = [], []    for row, query_id in zip(scores, query_ids):        ranked = [engine.docs[i]["id"] for i in np.argsort(-row)[:k]]        ndcgs.append(ndcg_at_k(ranked, qrels[query_id], k))        recalls.append(recall_at_k(ranked, qrels[query_id], k))    return float(np.mean(ndcgs)), float(np.mean(recalls))def main():    engine = HybridSearch()    data = json.loads((INDEX / "eval.json").read_text(encoding="utf-8"))    qrels = data["qrels"]    query_ids = sorted(qrels)    queries = [data["queries"][q] for q in query_ids]    print(f"queries {len(query_ids)}  documents {len(engine.docs)}")    started = time.perf_counter()    bm25 = np.array([engine.bm25_scores(q) for q in queries])    dense = np.array([engine.dense_scores(q) for q in queries])    print(f"scored both arms in {time.perf_counter() - started:.0f}s", file=sys.stderr)    arms = {        "bm25": bm25,        "dense": dense,        "rrf k=60": np.array([rrf(b, d) for b, d in zip(bm25, dense)]),    }    for alpha in ALPHAS:        arms[f"weighted a={alpha}"] = np.array(            [weighted(b, d, alpha) for b, d in zip(bm25, dense)]        )    print(f"{'arm':<18}{'nDCG@10':>9}{'Recall@10':>11}")    for name, matrix in arms.items():        ndcg, recall = score_matrix(engine, matrix, query_ids, qrels)        print(f"{name:<18}{ndcg:>9.4f}{recall:>11.4f}")if __name__ == "__main__":    main()

Run it. Scoring both arms across 323 queries took 9 to 24 seconds on the machines here, and the script prints nothing between the first line and the finished table.

bash
python evaluate.py

Expected output.

text
queries 323  documents 3633arm                 nDCG@10  Recall@10bm25                 0.3225     0.1471dense                0.3387     0.1580rrf k=60             0.3530     0.1660weighted a=0.0       0.3225     0.1471weighted a=0.1       0.3386     0.1602weighted a=0.2       0.3459     0.1625weighted a=0.3       0.3505     0.1635weighted a=0.4       0.3592     0.1674weighted a=0.5       0.3642     0.1712weighted a=0.6       0.3674     0.1716weighted a=0.7       0.3711     0.1737weighted a=0.8       0.3673     0.1707weighted a=0.9       0.3592     0.1672weighted a=1.0       0.3387     0.1580

What just happened. Four facts, and the rest of the service rests on them.

Hybrid genuinely beats both arms here: 0.3711 against 0.3225 for BM25 alone, a relative gain of about 15 percent, and against 0.3387 for dense alone. The sweep doubles as its own sanity check, because alpha=0.0 reproduces BM25 exactly and alpha=1.0 reproduces dense exactly. If your two endpoints do not match your single-arm rows, your fusion code has a bug.

Tuned weighting beats untuned RRF, 0.3711 against 0.3530. That is the argument for tuning at all.

And the curve has one hump, peaking at 0.7 instead of jumping around. A noisy curve would mean the corpus is too small to tune on, and the answer would be to take RRF and move on.

Step 8: Validate the fusion weight with cross-validation

Goal. Choose alpha with 5-fold cross-validation instead of reading the peak off the table you just printed.

Why this step. The table above tuned and reported on the same 323 queries, so 0.3711 is the best number this corpus can be made to show, not the number you should expect on new queries. The honest version tunes alpha on four fifths of the queries and reports on the fifth it never saw, five times over. If the tuned value is real, every fold picks it and the held-out score stays close to the tuned score. If it was noise, folds disagree and the held-out mean drops.

Code. Add this to the end of main() in evaluate.py:

python
    folds = np.array_split(np.random.default_rng(0).permutation(len(query_ids)), 5)    held_out = []    for fold in folds:        train = np.setdiff1d(np.arange(len(query_ids)), fold)        best = max(            ALPHAS,            key=lambda a: score_matrix(                engine,                np.array([weighted(bm25[i], dense[i], a) for i in train]),                [query_ids[i] for i in train],                qrels,            )[0],        )        ndcg, _ = score_matrix(            engine,            np.array([weighted(bm25[i], dense[i], best) for i in fold]),            [query_ids[i] for i in fold],            qrels,        )        held_out.append((best, ndcg))        print(f"fold: tuned alpha {best}  held-out nDCG@10 {ndcg:.4f}")    print(f"cross-validated nDCG@10 {np.mean([n for _, n in held_out]):.4f}")

Run it.

bash
python evaluate.py

Expected output (the arm table prints first, unchanged, then):

text
fold: tuned alpha 0.7  held-out nDCG@10 0.3174fold: tuned alpha 0.7  held-out nDCG@10 0.3947fold: tuned alpha 0.7  held-out nDCG@10 0.4365fold: tuned alpha 0.7  held-out nDCG@10 0.3683fold: tuned alpha 0.7  held-out nDCG@10 0.3381cross-validated nDCG@10 0.3710

What just happened. All five folds picked 0.7 independently, and the held-out mean of 0.3710 matches the tuned 0.3711 almost exactly. That is as clean a result as this procedure produces: alpha=0.7 is a property of this corpus, not an artifact of tuning.

Look at the spread, though: individual folds range from 0.3174 to 0.4365. Some groups of 65 queries are much easier than others. That spread is why a single score from a single query set is worth so little, and why a 0.005 difference between two configurations usually means nothing.

Step 9: Add the cross-encoder, then check whether it pays

Goal. Rerank the top candidates with a cross-encoder and measure what it buys.

Why this step. If you have not met cross-encoders before, how they score a query and a document together is worth ten minutes first. Reranking is the default advice for improving a retrieval pipeline, and it is usually given without a measurement. It has a real cost, paid per query at serving time rather than once at index time, so it needs a real benefit to justify it. You already have the harness to check.

Code. In search.py, change the reranker line in __init__:

python
        # max_length caps the query and document pair together. Truncating to 256        # halves the rerank latency on CPU and costs nothing measurable here.        self.reranker = CrossEncoder(RERANKER, max_length=256) if load_reranker else None

Then, in the search method, replace the line rerank_scores = {} with:

python
        rerank_scores = {}        if rerank and self.reranker is not None:            pairs = [(query, self.docs[i]["text"]) for i in candidates]            scores = self.reranker.predict(pairs, show_progress_bar=False)            rerank_scores = {i: float(s) for i, s in zip(candidates, scores)}            candidates.sort(key=lambda i: rerank_scores[i], reverse=True)

In evaluate.py, make the engine load the reranker when asked, by changing the first line of main():

python
    engine = HybridSearch(load_reranker="--rerank" in sys.argv)

and add this after the cross-validation block you added in step 8, at the end of main():

python
    if engine.reranker is not None:        alpha = held_out[0][0]        fused = np.array([weighted(b, d, alpha) for b, d in zip(bm25, dense)])        for depth in (10,):            started = time.perf_counter()            reranked = np.full_like(fused, -1e9)            for i, query in enumerate(queries):                top = np.argsort(-fused[i])[:depth]                pairs = [(query, engine.docs[j]["text"]) for j in top]                reranked[i, top] = engine.reranker.predict(pairs, show_progress_bar=False)            ndcg, recall = score_matrix(engine, reranked, query_ids, qrels)            per_query = (time.perf_counter() - started) / len(queries)            print(f"rerank top-{depth:<11}{ndcg:>9.4f}{recall:>11.4f}")            print(f"rerank cost {per_query:.2f}s per query", file=sys.stderr)

Documents that were never candidates get a score of negative 1e9, which keeps them below every reranked document while leaving the array shape intact for the scorer.

Run it. First check the search.py edit on its own, because the evaluation below calls the reranker directly and never goes through search, so a mistake in that branch would not show up here:

bash
python -c "from search import HybridSearch; \e = HybridSearch(load_reranker=True); \q = 'does vitamin d reduce cancer risk'; \hit = e.search(q, top_k=2, rerank=10)[0]; \print(hit['doc_id'], hit['fused'], hit['rerank'])"
text
MED-4570 0.9335 8.3613

A rerank of None there means the branch never ran, and a KeyError means the sort key is wrong. Then run the measurement:

bash
python evaluate.py --rerank

This downloads the cross-encoder on first use, about 90 MB, then reranks 10 candidates for each of the 323 queries. That took between 5 and 9 minutes across runs here.

Expected output (the last line, after the table and the folds):

text
rerank top-10            0.3690     0.1737

and on stderr, the cost, which ranged from rerank cost 0.84s per query to 1.51s per query across five runs on three machines.

What just happened. The cross-encoder scored 0.3690, which is 0.002 below the 0.3711 hybrid it reranked, and it cost between 0.84 and 1.51 seconds per query across five runs on three quiet machines to do it. A loaded laptop doubled that. Recall is identical by construction, because reranking only reorders candidates the fusion step already selected.

So on this corpus, at this candidate depth, reranking costs well over a second per query and buys nothing. That is a real result, not a failed experiment, and it is worth being precise about what it does and does not say. It does not say cross-encoders are useless. ms-marco-MiniLM-L6-v2 was trained on short web search queries against short passages; NFCorpus asks medical questions against abstracts, and the fusion step has already done well. A larger reranker such as bge-reranker-base, which has twelve times the parameters, may well do better, and on a corpus where the first stage is weaker there is far more room to improve.

What it does say is that you cannot know without measuring, and that the measurement takes one afternoon. The service therefore ships with reranking off by default, behind an environment variable, so that turning it on is a decision someone makes deliberately after running this script on their own data.

Step 10: Serve it with FastAPI

Goal. Put the engine behind an HTTP endpoint that returns every stage's score.

Why this step. Two things in this file are easy to get wrong and expensive to debug. The models load once at startup through a lifespan handler, not per request, because loading them takes seconds. And the endpoint is a plain def, not async def, because encoding a query is CPU-bound: FastAPI runs a plain def endpoint in a threadpool, while a blocking call inside an async def stalls the entire event loop for every other request.

Code. Create app.py:

python
"""FastAPI service exposing every stage's score for one query."""import osfrom contextlib import asynccontextmanagerfrom typing import Annotatedfrom fastapi import FastAPI, Queryfrom fastapi.responses import FileResponsefrom search import HybridSearchRERANK_ENABLED = os.environ.get("RERANK", "0") == "1"engine: HybridSearch | None = None@asynccontextmanagerasync def lifespan(app: FastAPI):    global engine    engine = HybridSearch(load_reranker=RERANK_ENABLED)    yield    engine = Noneapp = FastAPI(title="Hybrid search", lifespan=lifespan)@app.get("/")def home():    return FileResponse("static/index.html")@app.get("/health")def health():    return {"status": "ok", "documents": len(engine.docs), "rerank": RERANK_ENABLED}@app.get("/search")def search(    q: Annotated[str, Query(min_length=2, max_length=400)],    alpha: Annotated[float, Query(ge=0.0, le=1.0)] = 0.7,    fusion: Annotated[str, Query(pattern="^(weighted|rrf)$")] = "weighted",    top_k: Annotated[int, Query(ge=1, le=20)] = 5,    rerank: Annotated[int, Query(ge=0, le=30)] = 0,):    """Plain def, not async def: FastAPI runs it in a threadpool, so the    CPU-bound encoder call does not block the event loop."""    results = engine.search(        q, alpha=alpha, fusion=fusion, top_k=top_k, rerank=rerank if RERANK_ENABLED else 0    )    return {"query": q, "fusion": fusion, "alpha": alpha, "results": results}

/ serves the page you build in step 11, so it errors until then. The checks below use /health and /search instead.

The lifespan handler is the current way to run startup code. The older @app.on_event("startup") decorator is deprecated, and Starlette 1.0 removed its own implementation of it in March 2026, which broke containers that pinned an old FastAPI against a new Starlette.

Run it.

bash
python -m uvicorn app:app --port 8000

In a second terminal:

bash
curl -s "localhost:8000/health"

Expected output.

text
{"status":"ok","documents":3633,"rerank":false}

Now try an invalid parameter, because parameter validation you never tested is parameter validation you do not have:

bash
curl -s "localhost:8000/search?q=vitamin&alpha=2" | python -m json.tool

Expected output (the first lines of a 422 response):

json
{    "detail": [        {            "type": "less_than_equal",            "loc": [                "query",                "alpha"            ],            "msg": "Input should be less than or equal to 1",            "input": "2"

What just happened. The service is up, it loaded the index once, and it rejects an out-of-range alpha with a message that names both the parameter and the constraint. The rerank field in /health reports what the service is actually configured to do. Check that field first when someone reports that the API is slow.

Step 11: Add the browser page with an alpha slider

Goal. Give the service a page where you can move alpha and watch the ranking change.

Why this step. The evaluation script tells you which alpha is best on average across 323 queries. It cannot tell you what changed for the query a colleague is complaining about. A slider and a table of per-stage scores answer that in seconds, and the whole thing is one file with no build step.

Code. Create static/index.html:

html
<!doctype html><title>Hybrid search</title><style>  body { font: 15px system-ui, sans-serif; margin: 2rem auto; max-width: 52rem; }  input[type=search] { width: 100%; padding: .5rem; font-size: 1rem; }  table { border-collapse: collapse; width: 100%; margin-top: 1rem; }  th, td { border-bottom: 1px solid #ddd; padding: .4rem; text-align: left; }  td.num { text-align: right; font-variant-numeric: tabular-nums; }  label { margin-right: 1rem; }</style><h1>Hybrid search</h1><input type="search" id="q" placeholder="ask something medical" autofocus><p>  <label>alpha <input type="range" id="alpha" min="0" max="1" step="0.1" value="0.7">    <output id="alphaOut">0.7</output></label>  <label><input type="radio" name="fusion" value="weighted" checked> weighted</label>  <label><input type="radio" name="fusion" value="rrf"> RRF</label>  <label><input type="checkbox" id="rerank"> rerank top 10</label></p><table>  <thead><tr>    <th>title</th><th>BM25</th><th>dense</th><th>fused</th><th>rerank</th>  </tr></thead>  <tbody id="rows"></tbody></table><script>  const $ = (id) => document.getElementById(id);  const run = async () => {    const query = $("q").value.trim();    $("alphaOut").value = $("alpha").value;    if (query.length < 2) { $("rows").innerHTML = ""; return; }    const params = new URLSearchParams({      q: query,      alpha: $("alpha").value,      fusion: document.querySelector("input[name=fusion]:checked").value,      rerank: $("rerank").checked ? 10 : 0,    });    const data = await (await fetch("/search?" + params)).json();    $("rows").innerHTML = data.results.map((r) => `<tr>      <td>${r.title}</td>      <td class="num">${r.bm25}</td>      <td class="num">${r.dense}</td>      <td class="num">${r.fused}</td>      <td class="num">${r.rerank === null ? "-" : r.rerank}</td></tr>`).join("");  };  ["q", "alpha", "rerank"].forEach((id) => $(id).addEventListener("input", run));  document.querySelectorAll("input[name=fusion]").forEach(    (el) => el.addEventListener("change", run));</script>

Run it. Restart uvicorn, then open http://localhost:8000 and search for does vitamin d reduce cancer risk. Drag the slider from 0.7 to 0.0.

Expected output. The table reorders as you drag. At alpha=0.0 the ranking is pure BM25 and the dense column stops predicting the order; at alpha=1.0 the reverse. To check the same thing from the command line:

bash
curl -s "localhost:8000/search?q=does+vitamin+d+reduce+cancer+risk&fusion=rrf&top_k=2"

The first hit stays MED-4570, but its fused score is now 0.0312 rather than 0.9335, because RRF scores live on a different scale entirely.

What just happened. You can now answer "why is this document ranked here" for any single query. An evaluation table cannot answer that, and it is the question people actually bring you.

Step 12: Containerize the search service with Docker

Goal. Build an image that starts without downloading anything, and run it.

Why this step. The default torch wheel on PyPI for Linux pulls CUDA packages that are useless in a CPU image and cost gigabytes, so the build installs from PyTorch's CPU index first. The other trap is a container that downloads its model on startup. It fails whenever the network is slow, blocked or simply absent. Bake the weights in at build time, then set HF_HUB_OFFLINE=1 so a missing cache fails loudly during the build rather than quietly on the first cold start in production.

Code. Create requirements-serve.txt, which is the runtime subset. ir_datasets is only needed for building the index, so it stays out of the image:

text
bm25s==0.3.11PyStemmer==3.1.0sentence-transformers==6.1.0numpy==2.5.3fastapi==0.141.1uvicorn==0.53.0

Create .dockerignore:

text
__pycache__/.venv/

Create Dockerfile:

dockerfile
FROM python:3.13-slimENV HF_HOME=/models \    PYTHONUNBUFFERED=1WORKDIR /app# CPU-only torch. The default PyPI wheel for Linux pulls CUDA packages you# cannot use in this image and that cost gigabytes.RUN pip install --no-cache-dir --index-url https://download.pytorch.org/whl/cpu \    torch==2.14.0COPY requirements-serve.txt .RUN pip install --no-cache-dir -r requirements-serve.txt# Bake both models into the image so the container never downloads at startup,# then forbid Hub calls at runtime. The reranker is baked even though it is off# by default: RERANK=1 with no cached weights and no network cannot start.RUN python -c "\from sentence_transformers import SentenceTransformer, CrossEncoder; \SentenceTransformer('BAAI/bge-small-en-v1.5'); \CrossEncoder('cross-encoder/ms-marco-MiniLM-L6-v2')"ENV HF_HUB_OFFLINE=1COPY search.py app.py ./COPY static ./staticCOPY index ./indexEXPOSE 8000CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]

--host 0.0.0.0 is not optional. Uvicorn binds to 127.0.0.1 by default, which inside a container means the container's own loopback, and your curl on the host gets a connection reset with no error in the logs.

Run it.

bash
docker build -t hybrid-search .docker run -d --name hs -p 8000:8000 hybrid-searchcurl -s "localhost:8000/health"

From a fully pruned builder cache the build took 543 and 609 seconds on two runs here, most of it downloading torch. When only the model-baking layer changed it took 141 to 265 seconds, and when nothing changed, 3.5 seconds. It produces a 2.4 GB image.

Expected output.

text
{"status":"ok","documents":3633,"rerank":false}

And the same query as step 6, now served from the container:

bash
curl -s "localhost:8000/search?q=does+vitamin+d+reduce+cancer+risk&top_k=2"

returns MED-4570 with "bm25": 4.3297, "dense": 0.8108, "fused": 0.9335, identical to the local run. That is what you want from a container: a boundary, and no change in behaviour.

What just happened. You have a self-contained image with the index and the model weights inside it. docker logs hs shows the startup sequence ending in:

text
INFO:     Application startup complete.INFO:     Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit)

To try reranking in the container, pass the environment variable:

bash
docker rm -f hsdocker run -d --name hs -p 8000:8000 -e RERANK=1 hybrid-searchcurl -s "localhost:8000/health"
text
{"status":"ok","documents":3633,"rerank":true}

Now ask for reranked results:

bash
curl -s "localhost:8000/search?q=does+vitamin+d+reduce+cancer+risk&top_k=2&rerank=10" \  | python -m json.tool | cut -c1-88

The field that was null before now carries the cross-encoder's score:

text
        {            "doc_id": "MED-4570",            "title": "Vitamin D supplement doses and serum 25-hydroxyvitamin D in the ra            "bm25": 4.3297,            "dense": 0.8108,            "fused": 0.9335,            "rerank": 8.3613        },

That request took between 1.5 and 4.3 seconds end to end across machines here. It is the cost you measured in step 9, now paid per request.

This is also why the Dockerfile bakes the cross-encoder even though reranking is off by default. Bake only the embedder, and RERANK=1 starts a container that cannot reach Hugging Face because you just set HF_HUB_OFFLINE=1, so it exits with the OSError in the next section instead of serving.

Stop and remove it with docker rm -f hs.

When it breaks: the errors this build actually produces

Every error below came out of a real run of this build, not a list of things that could theoretically go wrong.

ModuleNotFoundError: No module named 'Stemmer'

Cause. The package is called PyStemmer, but it imports as Stemmer.

Fix. Install PyStemmer==3.1.0. The import line stays import Stemmer.

TypeError: Router.init() got an unexpected keyword argument 'on_startup'

Cause. An old FastAPI paired with Starlette 1.0 or newer, which removed startup events in March 2026.

Fix. Use the lifespan handler from step 10 and pin fastapi==0.141.1. Do not reach for @app.on_event("startup"), which is what this error is usually chasing.

OSError: We couldn't connect to 'https://huggingface.co'

Cause. HF_HUB_OFFLINE=1 with no cached weights, or no network on the first run. In this build it means the image set the offline flag but never baked in the model the container then asked for.

Fix. Confirm the RUN python -c ... download step ran during the build, and that HF_HOME is the same at build time and run time. If you enable reranking, the cross-encoder has to be baked in too.

curl: (56) Recv failure: Connection reset by peer

Cause. Uvicorn bound to 127.0.0.1 inside the container, which is the container's own loopback, not one your host can reach.

Fix. Add --host 0.0.0.0 to the CMD, as the Dockerfile in step 12 does.

ValueError: k of 5000 is larger than the number of available scores

Cause. Asking BM25 for more results than the corpus holds.

Fix. Cap the requested depth at the corpus size. The API caps top_k at 20 for this reason.

Fusion scores that all come out identical

Cause. A query whose terms are all outside the vocabulary, so BM25 returns a constant array and min-max normalisation has nothing to spread.

Fix. Nothing: this is the minmax guard from step 5 doing its job. Check the query rather than the code.

Every alpha gives the same score in the sweep

Cause. The query tokenisation does not match the index tokenisation, so the BM25 arm scores zero everywhere and only the dense arm moves.

Fix. Use the same stopwords setting and the same stemmer in bm25_scores as build_index.py used when indexing.

How the finished pieces fit together

mermaid
flowchart TD
    Q["Query string"] --> BM["BM25 arm<br/>bm25s + stemmer"]
    Q --> DE["Dense arm<br/>bge-small-en-v1.5"]
    BM --> BS["Score vector<br/>3633 unbounded scores"]
    DE --> DS["Score vector<br/>3633 cosine scores"]
    BS --> FU{"Fusion<br/>weighted or RRF"}
    DS --> FU
    FU --> TOP["Top candidates<br/>with per-stage scores"]
    TOP --> RR["Cross-encoder rerank<br/>off by default"]
    RR --> API["/search response"]
    TOP --> API
    IDX[("index/<br/>bm25, embeddings.npy,<br/>docs.jsonl")] --> BM
    IDX --> DE
    EV["evaluate.py<br/>sets alpha = 0.7"] -.-> FU

    style Q fill:#4A90E2,color:#FFFFFF
    style BM fill:#98D8C8,color:#2C2C2A
    style DE fill:#98D8C8,color:#2C2C2A
    style BS fill:#95A5A6,color:#FFFFFF
    style DS fill:#95A5A6,color:#FFFFFF
    style FU fill:#7B68EE,color:#FFFFFF
    style TOP fill:#6BCF7F,color:#2C2C2A
    style RR fill:#FFA07A,color:#2C2C2A
    style API fill:#4A90E2,color:#FFFFFF
    style IDX fill:#FFD93D,color:#2C2C2A
    style EV fill:#C2185B,color:#FFFFFF

What makes this service debuggable is visible in the diagram. The per-stage scores survive all the way to the response, rather than being collapsed into one number at the fusion step. And evaluate.py sits outside the request path, feeding it a single value: the alpha that step 8 showed holds up on queries it was not tuned on. Reranking sits off to the side, skipped unless RERANK=1. That is why the response carries a rerank field that is usually null.

The complete artifact

text
hybrid-search/├── .dockerignore├── Dockerfile├── requirements.txt          # build + serve├── requirements-serve.txt    # serve only, no ir_datasets├── explore.py                # step 1 only: prints what the corpus contains├── build_index.py            # one-time: downloads NFCorpus, writes index/├── search.py                 # HybridSearch, minmax, ranks_from, rrf, weighted├── evaluate.py               # arm table, alpha sweep, cross-validation, rerank├── app.py                    # FastAPI service├── static/│   └── index.html            # slider UI└── index/                    # built artifacts, about 15 MB    ├── bm25/    ├── docs.jsonl    ├── embeddings.npy    └── eval.json

Here is search.py in its final state, with every edit from steps 4, 5, 6 and 9 merged in:

python
"""Hybrid search over the prebuilt indexes: BM25, dense, fusion, optional rerank."""import jsonimport pathlibimport bm25simport numpy as npimport Stemmerfrom sentence_transformers import CrossEncoder, SentenceTransformerINDEX = pathlib.Path("index")EMBEDDER = "BAAI/bge-small-en-v1.5"RERANKER = "cross-encoder/ms-marco-MiniLM-L6-v2"def minmax(scores):    """Rescale one query's scores to [0, 1]. Constant input maps to all zeros."""    low, high = float(scores.min()), float(scores.max())    if high <= low:        return np.zeros_like(scores)    return (scores - low) / (high - low)def ranks_from(scores):    """1-based rank of every document: the top-scoring document gets rank 1."""    return np.argsort(np.argsort(-scores)) + 1def rrf(bm25_scores, dense_scores, k=60):    """Reciprocal rank fusion (Cormack et al., SIGIR 2009)."""    return 1.0 / (k + ranks_from(bm25_scores)) + 1.0 / (k + ranks_from(dense_scores))def weighted(bm25_scores, dense_scores, alpha):    """Convex combination of min-max normalised scores. alpha weights the dense arm."""    return alpha * minmax(dense_scores) + (1.0 - alpha) * minmax(bm25_scores)class HybridSearch:    def __init__(self, index_dir=INDEX, load_reranker=False):        index_dir = pathlib.Path(index_dir)        lines = (index_dir / "docs.jsonl").read_text(encoding="utf-8")        self.docs = [json.loads(line) for line in lines.splitlines()]        self.bm25 = bm25s.BM25.load(str(index_dir / "bm25"), show_progress=False)        self.embeddings = np.load(index_dir / "embeddings.npy")        self.stemmer = Stemmer.Stemmer("english")        self.embedder = SentenceTransformer(EMBEDDER)        # max_length caps the query and document pair together. Truncating to 256        # halves the rerank latency on CPU and costs nothing measurable here.        self.reranker = CrossEncoder(RERANKER, max_length=256) if load_reranker else None    def bm25_scores(self, query):        tokens = bm25s.tokenize(            query,            stopwords="en",            stemmer=self.stemmer,            return_ids=False,            show_progress=False,        )        return self.bm25.get_scores(tokens[0])    def dense_scores(self, query):        vector = self.embedder.encode_query(query, normalize_embeddings=True)        return self.embeddings @ vector    def search(self, query, alpha=0.7, fusion="weighted", top_k=5, rerank=0):        bm25 = self.bm25_scores(query)        dense = self.dense_scores(query)        fused = rrf(bm25, dense) if fusion == "rrf" else weighted(bm25, dense, alpha)        depth = max(top_k, rerank)        candidates = list(np.argsort(-fused)[:depth])        rerank_scores = {}        if rerank and self.reranker is not None:            pairs = [(query, self.docs[i]["text"]) for i in candidates]            scores = self.reranker.predict(pairs, show_progress_bar=False)            rerank_scores = {i: float(s) for i, s in zip(candidates, scores)}            candidates.sort(key=lambda i: rerank_scores[i], reverse=True)        results = []        for position in candidates[:top_k]:            results.append(                {                    "doc_id": self.docs[position]["id"],                    "title": self.docs[position]["title"],                    "bm25": round(float(bm25[position]), 4),                    "dense": round(float(dense[position]), 4),                    "fused": round(float(fused[position]), 4),                    "rerank": round(rerank_scores[position], 4)                    if position in rerank_scores                    else None,                }            )        return resultsif __name__ == "__main__":    engine = HybridSearch()    for hit in engine.search("does vitamin d reduce cancer risk", top_k=3):        print(            f"{hit['doc_id']:<9} bm25 {hit['bm25']:>7.4f}  dense {hit['dense']:>6.4f}"            f"  fused {hit['fused']:>6.4f}  {hit['title'][:34]}"        )

evaluate.py is the other file assembled across several steps: created in step 7, then extended at the end of main() in step 8 and again in step 9. Here is that function in its final state, so you can diff against it if your numbers disagree with the page:

python
def main():    engine = HybridSearch(load_reranker="--rerank" in sys.argv)    data = json.loads((INDEX / "eval.json").read_text(encoding="utf-8"))    qrels = data["qrels"]    query_ids = sorted(qrels)    queries = [data["queries"][q] for q in query_ids]    print(f"queries {len(query_ids)}  documents {len(engine.docs)}")    started = time.perf_counter()    bm25 = np.array([engine.bm25_scores(q) for q in queries])    dense = np.array([engine.dense_scores(q) for q in queries])    print(f"scored both arms in {time.perf_counter() - started:.0f}s", file=sys.stderr)    arms = {        "bm25": bm25,        "dense": dense,        "rrf k=60": np.array([rrf(b, d) for b, d in zip(bm25, dense)]),    }    for alpha in ALPHAS:        arms[f"weighted a={alpha}"] = np.array(            [weighted(b, d, alpha) for b, d in zip(bm25, dense)]        )    print(f"{'arm':<18}{'nDCG@10':>9}{'Recall@10':>11}")    for name, matrix in arms.items():        ndcg, recall = score_matrix(engine, matrix, query_ids, qrels)        print(f"{name:<18}{ndcg:>9.4f}{recall:>11.4f}")    folds = np.array_split(np.random.default_rng(0).permutation(len(query_ids)), 5)    held_out = []    for fold in folds:        train = np.setdiff1d(np.arange(len(query_ids)), fold)        best = max(            ALPHAS,            key=lambda a: score_matrix(                engine,                np.array([weighted(bm25[i], dense[i], a) for i in train]),                [query_ids[i] for i in train],                qrels,            )[0],        )        ndcg, _ = score_matrix(            engine,            np.array([weighted(bm25[i], dense[i], best) for i in fold]),            [query_ids[i] for i in fold],            qrels,        )        held_out.append((best, ndcg))        print(f"fold: tuned alpha {best}  held-out nDCG@10 {ndcg:.4f}")    print(f"cross-validated nDCG@10 {np.mean([n for _, n in held_out]):.4f}")    if engine.reranker is not None:        alpha = held_out[0][0]        fused = np.array([weighted(b, d, alpha) for b, d in zip(bm25, dense)])        for depth in (10,):            started = time.perf_counter()            reranked = np.full_like(fused, -1e9)            for i, query in enumerate(queries):                top = np.argsort(-fused[i])[:depth]                pairs = [(query, engine.docs[j]["text"]) for j in top]                reranked[i, top] = engine.reranker.predict(pairs, show_progress_bar=False)            ndcg, recall = score_matrix(engine, reranked, query_ids, qrels)            per_query = (time.perf_counter() - started) / len(queries)            print(f"rerank top-{depth:<11}{ndcg:>9.4f}{recall:>11.4f}")            print(f"rerank cost {per_query:.2f}s per query", file=sys.stderr)

Its imports, metric functions and __main__ block are unchanged from step 7.

build_index.py, app.py, static/index.html and the Dockerfile are complete as written in steps 2, 3, 10, 11 and 12; nothing was left out of them.

Where to go next

Run the evaluation on your own corpus. This is the only step that changes what you ship. You need queries and judgments, and you need fewer than you think: evaluate.py works on any corpus you can put into the same three files. Fifty queries with a handful of judged documents each is enough to tell 0.3 from 0.7, though step 8's fold spread shows it is not enough to tell 0.68 from 0.72.

Try a stronger reranker on the same harness. Swap RERANKER for BAAI/bge-reranker-base and re-run evaluate.py --rerank. It has about twelve times the parameters, so expect it to be several times slower per query on CPU. The question to answer is whether it clears the 0.3711 bar that the fusion step already reaches. If no pretrained reranker does, fine-tune the reranker on your own labeled pairs and run the same check.

Move the index into a vector database and compare. Qdrant, Elasticsearch, OpenSearch and Weaviate all now do hybrid retrieval natively, and their defaults differ in ways that matter: Qdrant's RRF uses k=2 over zero-based ranks, Elasticsearch and OpenSearch use 60 over one-based ranks, and Weaviate's alpha follows the same convention as this tutorial's, where 1 means pure vector search. Port the corpus, run the same evaluation against the engine's own fusion, and you will know whether its defaults suit your data or just its benchmarks.

Both companion tutorials go one layer down on what this page deploys: the bake-off harness compares five retrievers on a different BEIR dataset, and reranking with a cross-encoder digs into the stage this build measured and then switched off.

References


AI Engineering

Information Retrieval

Follow for more technical deep dives on AI/ML systems, production engineering, and building real-world applications:


Books by Ranjan Kumar

Harness Engineering for Production AI Systems cover

Harness Engineering

The 7 GenAI Architectures cover

The 7 GenAI Architectures

Building Real-World Agentic AI Systems with LangGraph cover

Building Real-World Agentic AI Systems

The ChatML Handbook cover

The ChatML Handbook

The Chat Templates Handbook cover

The Chat Templates Handbook

Comments