Yandex Puts BM25 and Vectors in One SQL Query — and Reminds Everyone Why Exact Match Still Matters

Yandex’s cloud arm has shipped something that sounds like plumbing but reads like a retrieval lesson for every SEO worrying about AI answers. On 2 October, Yandex B2B Tech announced hybrid search in its YDB database: full-text search with BM25 ranking and vector (semantic) search, combined inside a single SQL query. CNews carried the company’s claim that YDB is the first Russian database with built-in hybrid search; the feature is in the on-premises enterprise edition of YDB 26.3 and in Managed Service for YDB.

To be clear up front: this is a B2B database product, not a Yandex Search ranking change. Nothing in the announcement says Alice or the main SERP runs on it. But the engineering write-up explains, unusually candidly, why semantic search on its own keeps missing the things users actually type.

What shipped

The detail is in the YDB team’s long Habr post by developer Alexander Zevaykin. The headline points:

  • Full-text indexes with ranking. FulltextMatch selects documents containing the query terms; FulltextScore ranks them with classic BM25 — term frequency, document length and term rarity.
  • Hybrid ranking in one query. A new HybridRank function takes a full-text branch and a vector branch and fuses them. The default is Reciprocal Rank Fusion with a constant of 60; a “linear” mode normalises scores and applies weights you choose.
  • Indexes live inside the database. Both index types are YDB service tables, updated in the same transaction as the data, so a deleted or out-of-stock row disappears from results immediately rather than after an ETL lag.
  • Yandex morphology for paying customers. The open-source build stems with Snowball; Yandex Cloud and enterprise customers get “Yandex morphology, the same that works in Yandex search”.

The post says full-text and hybrid search are in open source and have been available in the managed service, including Serverless, since 23 September. YDB CTO Andrey Fomichev told Habr the main engineering challenge was merging and ranking the two result sets.

Why SEOs should read the Habr post anyway

The examples are the gift. A lawyer searches for payouts under policy “7702-345678” after a flight change: the policy number must match character for character, while the incident description needs semantic matching. The team notes that vector embeddings do not store exact character sequences, so a rare surname or a unique number can fall outside the top 50 “similar” documents — and that Russian federal laws “152-ФЗ” and “153-ФЗ” sit right next to each other in vector space despite being different laws. Full-text search fails the other way: “lease agreement” will never match “rental contract” on words alone.

That is exactly the failure pattern publishers see in AI answers. Retrieval-augmented systems pick a few dozen passages before the model writes a word. If your product page says “our flagship model” instead of the actual model number, a semantic retriever may still find you — but so will it find five competitors, and the exact-match signal that should have broken the tie is missing from your copy.

The second lesson is the candidate pool. In YDB each branch gathers a pool of candidates — by default ten times the query LIMIT — before fusion. The authors are blunt: a document that does not make its branch’s pool cannot be rescued by any amount of clever merging. Swap “document” for “your page” and “fusion” for “answer generation”, and you have the clearest public explanation of why citation tracking is a retrieval problem first and a writing problem second.

Our take

Yandex has form for publishing the guts of its search stack — see the AI-Blender’s 50-millisecond SERP decisions and the Alice AI Search pretrain released to researchers. This one is aimed at developers building marketplace search, support bots and RAG assistants, and that is precisely why it matters for CIS ecommerce: the next generation of on-site search and shopping assistants in Russia may well be built on it.

Practical takeaways: keep identifiers literal and visible — SKUs, model numbers, standards, error codes, place names — in body text, not just in images or JavaScript widgets. Write the plain-language description too, because the vector branch needs it. And stop treating “keywords versus semantics” as an either/or. The database people have quietly settled that argument: you need both, in the same query, at the same time.