Yandex’s Alice AI Search team has published a rare look under the bonnet of its generative quick answers — and open-sourced the base pretrain. In an 11 September 2026 Habr deep dive, the team describes Alice AI-T5-35B-A0.6B: a hybrid encoder-decoder Mixture-of-Experts model trained from scratch for fast, grounded answers on Yandex Search. Most regional engines wave at “AI search”; few publish architecture notes this concrete.
The interesting bit for search marketers is not the parameter count. It is the pipeline. Yandex runs agentic retrieval, then a tiny ~80M extractor that trims documents into “info-contexts” before generation. They claim that cut input tokens enough to lift throughput by about 40% without quality loss. Alignment mixes SFT, online RL (GSPO/GRPO), and behavioural “surplus” rewards from how people actually use Search. In plain English: the model is trained partly on whether users keep going, clarify, or bounce — not only on annotator preference scores.
Yandex says side-by-side tests beat Google AI Overviews in 56.7% of comparisons after the June release, with online A/B lifts on sessions-per-user. Vendor benchmarks again — treat them as directional, not gospel. But the architecture story is consistent with what we are seeing at Brave, Perplexity, and others: context quality and latency budgets beat raw frontier-model flex for search answers. If your pages are bloated, poorly chunkable, or bury the answer under partner widgets, extractors will skip you even when classic ranking still likes the URL.
External researchers can run inference via Hugging Face Transformers; Yandex’s production stack stays internal. For anyone tracking non-Google AI SEO, this is one of the most concrete public documents yet on how a major regional engine builds answer cards at scale. Pair it with the later 49.5 million monthly user claim for Alice quick answers and you get both the lab notebook and the traffic reality.
Caveats for practitioners: open-sourcing the pretrain does not mean you can reproduce production answers, and a win rate versus Google AI Overviews is not a substitute for logging when Alice cites you on your money queries. Do that weekly. Also watch whether “info-context” extraction privileges listicles, tables, and definitional openings over long narrative posts — a pattern that already shows up in other AI-SERP stacks.
There is a secondary lesson for content teams. If Yandex is rewarding compact info-contexts, your CIS pages should lead with the answerable claim, then evidence, then narrative — not the reverse. That is not “writing for AI” as a slogan; it is matching an extractor that has a latency budget. Teams that only optimise title tags will keep wondering why Alice cites a dull but dense competitor page.
Finally, treat the Hugging Face release as competitive intelligence, not a deployment plan. Few SEOs will fine-tune Alice AI-T5 locally — but reading what the model was trained to prefer tells you which page shapes survive the funnel into an answer card. Bookmark the Habr piece next to your Yandex Webmaster notes and revisit after each major Alice UI change.





Leave a Reply