Skip to content
AI & Automation

Hybrid Search for RAG: Combining Keyword and Semantic Search

How hybrid search works in RAG: BM25 keyword search, vector search, reciprocal rank fusion and weighted fusion, filters, sparse vectors, relevance tuning and implementation options.

Quick answer

Hybrid search runs keyword search (typically BM25) and vector search for the same query and merges the results. Keyword search catches exact terms such as product codes, names and error messages; vector search catches paraphrases and intent. Merge with reciprocal rank fusion, which combines rankings without comparing raw scores, or with tuned weighted score fusion, apply permission and metadata filters in both, then rerank the fused candidates before generation. Tune and verify the combination on your own questions.

Where This Fits

Hybrid search is the retrieval stage of RAG, usually followed by reranking. Storage options are compared in vector databases, and the same idea applied to product search is in ecommerce semantic search.

Keyword (BM25)Vector (semantic)
MatchesExact terms and their frequencyMeaning and similarity
Strong onCodes, IDs, names, rare words, quotesParaphrases, synonyms, natural questions
Weak onDifferent wording for same ideaExact identifiers, negation, numbers
InfrastructureInverted indexVector index
ExplainabilityVisible term matchesSimilarity scores

How Hybrid Retrieval Works

Fusion merges two imperfect rankings into a better candidate set for the reranker.

Fusion Methods

Reciprocal rank fusion (RRF) gives each document a score based on its rank in each list, roughly the sum of 1 divided by (constant plus rank), and orders by that. Because it ignores raw scores, it combines very different scoring systems robustly; Elasticsearch's RRF uses a rank constant that defaults to 60, and Qdrant supports RRF in hybrid queries.

Weighted score fusion normalizes scores from each method and combines them with weights. Weaviate's hybrid search exposes this through an alpha parameter, where 0 is pure keyword and 1 pure vector. It can outperform RRF when tuned, but tuning must be done on evaluation data.

Example: reciprocal rank fusion (pseudocode)
rrf(lists, k = 60):
  scores = {}
  for ranked in lists:                 # e.g. [bm25_results, vector_results]
    for rank, doc in enumerate(ranked, start = 1):
      scores[doc] += 1 / (k + rank)
  return sort_by_value_desc(scores)

Filters and Permissions

Apply the same filters to both retrievers: permissions, tenant, product, language and date. A filter applied to only one side can leak restricted content through the other. Check how your system applies filters with approximate vector indexes, which can return fewer results after filtering.

Your RAG system missing exact codes and names?

ZSpace Labs implements hybrid retrieval with fusion, filters and reranking, tuned on your own questions.

Start a Project

Implementation Options

OptionHow hybrid worksNotes
Elasticsearch / OpenSearchBM25 plus kNN with RRF or combined queriesMature keyword search
WeaviateBuilt-in hybrid with alpha weightingSingle query API
QdrantDense and sparse vectors with prefetch and fusionRRF and distribution-based fusion
PineconeSparse-dense vectorsManaged service
PostgresFull-text search plus pgvector, fused in SQL or codeKeeps data in one database

Tuning Relevance

  • Build a question set that includes identifier lookups and natural-language questions
  • Measure recall for keyword only, vector only and hybrid
  • Tune candidate counts per retriever and fusion parameters
  • Configure analyzers for your language, synonyms and code formats
  • Add reranking and measure again
  • Re-test when content or embedding models change

Advantages and Limitations

Hybrid search reliably improves recall across mixed query types and makes retrieval more robust. It adds two indexes to maintain, more parameters to tune and slightly more latency. For small, uniform corpora where one method already performs well, the extra complexity may not be needed.

How to Implement Step by Step

  • 1. Add a keyword index alongside the vector index, with the same chunks and metadata
  • 2. Run both retrievers with identical filters
  • 3. Fuse with RRF as a robust default
  • 4. Rerank the fused candidates
  • 5. Evaluate against single-method baselines
  • 6. Tune weights or fusion parameters only with evaluation data

Example: Hybrid Search in Postgres

Teams already on Postgres can combine built-in full-text search with pgvector and fuse results with reciprocal rank fusion in SQL. It is not as feature-rich as a search engine, but it keeps everything in one database.

Example: keyword plus vector retrieval fused with RRF (illustrative SQL)
WITH keyword AS (
  SELECT id, row_number() OVER (ORDER BY ts_rank_cd(tsv, q) DESC) AS rank
  FROM doc_chunks, websearch_to_tsquery('english', $1) q
  WHERE tsv @@ q AND tenant_id = $3
  ORDER BY ts_rank_cd(tsv, q) DESC LIMIT 50
),
semantic AS (
  SELECT id, row_number() OVER (ORDER BY embedding <=> $2) AS rank
  FROM doc_chunks
  WHERE tenant_id = $3
  ORDER BY embedding <=> $2 LIMIT 50
)
SELECT id, SUM(1.0 / (60 + rank)) AS rrf_score
FROM (SELECT * FROM keyword UNION ALL SELECT * FROM semantic) r
GROUP BY id
ORDER BY rrf_score DESC
LIMIT 20;

Which Retriever Wins for Which Query

Query typeUsually bestExample
Exact identifierKeyword'Error E-4021 on startup'
Natural questionVector'Why does my laptop lose VPN after sleep?'
Name or rare termKeyword'Halvorsen clause'
Concept with varied wordingVector'staff leave when a child is born'
MixedHybrid'refund policy for SKU 88-210 bought online'

Worked Example

An illustrative scenario, not a client case: an IT support assistant searches internal runbooks. Vector-only retrieval handles 'laptop won't connect to VPN' but misses queries quoting exact error codes. Adding BM25 and fusing with RRF makes error-code queries retrieve the right runbook, while natural-language questions perform as before.

Common Mistakes

  • Vector-only retrieval for identifier-heavy content
  • Different filters on the two retrievers
  • Combining raw scores without normalization
  • Tuning weights on a handful of queries
  • Skipping reranking after fusion

Need retrieval that handles every kind of question?

Talk to ZSpace Labs about hybrid search and RAG development.

Start a Project

Conclusion

Hybrid search combines the precision of keywords with the flexibility of meaning. Fuse rankings robustly, filter consistently, rerank and measure. Related: reranking, vector databases and RAG guide.

FAQ

Common questions

Search that combines keyword (lexical) retrieval such as BM25 with vector (semantic) retrieval and merges the results, so queries benefit from both exact term matching and meaning-based matching.

Related services
Relevant industries
Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.