I thought semantic search was the holy grail of retrieval until we launched our internal knowledge assistant. It looked perfect on paper, but in production it was a mess. One of our platform engineers asked a simple question about a message-queue retry policy, and the system gave a confident, three-paragraph answer about exponential backoff that was completely wrong. The document she actually needed was sitting at position eleven, just outside the top ten we passed to the LLM. That is why hybrid search and re-ranking isn’t an optional optimization for production systems, it’s a requirement.
For most developers, the go-to approach is “just embed it and use cosine similarity.” But dense retrieval, which converts text into high-dimensional vectors, has a big blind spot. It’s great at conceptual queries like “how do we handle incidents?” but it falls apart when someone searches for a specific technical term like “dead-letter queue threshold.” The embedding model averages those specific terms into a single vector and loses the precision you need for technical lookups.
Why BM25 still matters in the age of AI
Before we had neural retrieval, we had BM25 (Best Match 25). It’s a keyword-based approach that understands that rare terms matter more. If “dead-letter queue threshold” appears in only three documents, those documents get a big score boost. Unlike dense vectors, BM25 doesn’t care about semantic meaning, it cares about exact matches, which is what your engineers want when they search for specific config keys or error codes.
The solution isn’t to pick one, it’s to use both. We use hybrid search and re-ranking to combine their strengths. In systems like Weaviate or LlamaIndex, an alpha parameter blends the results: an alpha of 1.0 is pure vector search, and 0.0 is pure BM25. For technical documentation, the sweet spot is usually around 0.5.
If you’re wondering why your current setup keeps missing, take a look at my deep dive on why vector search misses exact IDs.
The two-stage funnel: re-ranking with cross-encoders
Even with a hybrid retriever, you still hit the “lost in the middle” problem. LLMs pay more attention to the start and end of their context window. If the most relevant chunk is ranked at #8 out of 10, the model might skip over it. This is where cross-encoders come in.
A bi-encoder (your standard embedding model) processes the query and document independently. A cross-encoder looks at them together. Because it sees the direct interaction between the question and the answer, it’s far more accurate, but also much slower. You can’t run it over a million documents, so you use it as a second stage to re-score the top 20 candidates your hybrid search returned.
from llama_index.postprocessor.sbert_rerank import SentenceTransformerRerank
from llama_index.core.query_engine import RetrieverQueryEngine
# Stage 2: Cross-encoder re-ranker (ms-marco is a solid general choice)
reranker_postprocessor = SentenceTransformerRerank(
model="cross-encoder/ms-marco-MiniLM-L-6-v2",
top_n=5
)
# Assemble the query engine with the re-ranker
query_engine = RetrieverQueryEngine.from_args(
retriever=retriever,
node_postprocessors=[reranker_postprocessor]
)
# Now the LLM only sees the 5 most statistically relevant chunks
response = query_engine.query("What is the retry limit for the payment service?")
Measuring success with RAGAS
Don’t just guess whether your search is better, measure it. Track Context Precision and Context Recall. When we added hybrid search and re-ranking, our Context Recall (finding the right doc at all) jumped from 0.74 to 0.83. Our Context Precision (how relevant the chunks we send to the LLM are) hit 0.79. That means fewer hallucinations and more honest “I don’t know” responses when the answer isn’t there.
Metadata filtering is another common trap. If you have a runbook for a service that was decommissioned two years ago, it shouldn’t even be in the candidate pool. Use filters to narrow the retrieval space by department or document age before the scoring starts. It’s faster and a lot more reliable.
If this hybrid search and re-ranking work is eating up your dev hours, let me handle it. I’ve been wrestling with WordPress since the 4.x days.
Stop guessing, start tuning
Production RAG works as a funnel. Start with a wide net using a hybrid retriever, narrow it with metadata filters, then use a cross-encoder to hand the best context to your LLM. If you’re still seeing confident wrong answers, your retrieval pipeline is probably the bottleneck, not your model choice. Fix the search logic first, and you’ll cut both LLM token costs and debugging time.