Sber GigaChat News

This feed includes only recent releases, roughly from the last 12 months, not the full historical archive. That is why a long-established product may have only a few news items.

2 news

Sber GigaChat
GigaChat 3.5 Ultra: hybrid linear attention architecture, weights are open

432B parameters (28B active, MoE) versus 702B in the previous Ultra - a hybrid of standard attention and linear attention (GatedDeltaNet): 4x less KV-cache per token, 2.14x more context in the same memory. Sber's first model to publish not only weights but also key training stages; claimed to outperform DeepSeek V4 Flash and V3.2 on coding.

By contrast, Western flagships (GPT, Gemini, Claude) are closed-weights and available only through API/cloud - GigaChat 3.5 Ultra publishes its weights OpenAI →
Sber GigaChat
GigaChat-3.1 Ultra and Lightning: Dense-to-MoE shift, new alignment

Ultra (702B MoE) and Lightning (10B, 1.8B active parameters) were moved from a dense architecture to MoE, and the post-training stage was redesigned (DPO in FP8: higher quality than bf16 with half the memory). The weights are open under MIT on HuggingFace and GitVerse; Ultra outperforms non-reasoning Qwen3-235B-A22B and DeepSeek-V3-0324 in math and general reasoning.

Ultra is directly compared with Qwen3-235B-A22B on math and reasoning Alibaba Qwen →

Discuss Sber GigaChat News

Send via: