Yandex Trains AliceAI-Foundation-80B From Scratch — and Open-Sources the Base

Yandex just drew a brighter line under its Alice AI stack: the company has released AliceAI-Foundation-80B-A3B-Base — an 80-billion-parameter Mixture-of-Experts base model with about 3B active parameters per token — trained entirely from scratch at Yandex and published under Apache 2.0 on Hugging Face. The engineering write-up on Yandex’s Medium channel is unusually frank about failed runs, bad scaling-law forecasts and why agentic pretrain data matters. For CIS SEO and anyone tracking who actually owns the models behind AI answers in Search, this is a different story from last week’s product unlocks.

Two earlier Alice releases still matter and should not be confused with this one. In September Yandex open-sourced the smaller Alice AI Search Pretrain used for fast answers under the search box. Separately, Alice chat gained automatic model routing and Expert mode, and Yandex later unlocked flagship chat models for free with live search citations. Foundation-80B is the research/base layer aimed at post-training and agentic capabilities — not another SERP widget announcement.

What the model actually is

Per the model card and Medium post, AliceAI-Foundation-80B-A3B-Base is a hybrid MoE Transformer: 48 blocks, hidden size 2048, 512 routed experts with top-k 10 plus a shared expert, multi-token prediction, and a claimed context window up to 262,144 tokens. Attention mixes Kimi Delta Attention with full attention in a 3:1 pattern, plus Attention Residuals. Training ran in stages from an 8K main pretrain through 32K and 256K context extension, then a final reasoning/agentic stage. Yandex stresses it did not initialise from third-party open weights this time — a deliberate break from the earlier Alice AI LLM line that started from Qwen3–235B-A22B.

On Yandex’s internal suite the new base leads on most Russian factual, education and expert slices (WikiWebFacts, HardMultiQA, CultCat, EduBench Russian/History/Literature, ExpertFactsQA medicine/law) while staying competitive on MATH-500, LiveCodeBench and long-context FinQA 128k against larger peers including DeepSeek-V4-Flash-Base (284B-A13B), Nemotron-3-Super and Qwen3.5–35B-A3B-Base. Treat vendor benches as vendor benches — but the comparison table is unusually wide for a CIS search company, and Yandex is also releasing two Russian factual benchmarks with evaluation protocols.

Why search people should care

Opinion: open-sourcing the base that will underpin Alice’s agentic path is a signal about who controls the evidence stack, not just another Hugging Face trophy. If Alice’s future tool-using answers are post-trained from a Yandex-owned pretrain rather than a thinly wrapped foreign base, the citation, safety and language behaviour of Yandex Search AI becomes harder for competitors to clone and easier for Yandex to tune against Russian query distributions. That is strategically closer to how Baidu thinks about ERNIE-in-Search than to a pure “rent Bing + slap a chat panel” European privacy engine.

The Medium post’s operational lessons are the part SEOs should steal even if they never download 80B weights. Held-out corpus loss was a poor predictor of benchmark winners. A scaling-law learning rate “blew up” MoE activations until they moved a Z-loss idea onto MoE outputs. Factual augmentations that worked in short fine-tunes had to be scaled across ~30% of useful web facts for long pretrain gains. Hard reasoning data alone did not create agent skills — tool-interaction trajectories in pretrain did. Strip the GPU romance and you still get a useful checklist for anyone building retrieval-augmented answer products: corpus coverage beats clever loss curves; agent data must appear before RL if you want stable tool use.

What not to overclaim

This is a base model. Yandex says it is for research and further training, not a drop-in chat product. It does not, by itself, change blue-link rankings tomorrow morning. It does sit next to Alice Search’s 49M+ monthly quick-answer users and the newer Expert harness as evidence that Yandex is consolidating Search, chat and agents around an in-house model family. If you publish for Russian search, keep watching Webmaster “Alice visibility” style reports and citation behaviour as those post-trained descendants ship — and keep linking your money pages to clear, dated, extractable claims. Models trained with agentic search trajectories will reward pages that survive tool use, not just classic TF-IDF relevance.

Bottom line: Yandex is no longer only fine-tuning someone else’s open weights for Alice. Foundation-80B is the from-scratch bet. Download it if you post-train; track it if you care who owns the brains behind Yandex’s AI SERPs.