Prices per 1 million tokens (input / output), unless otherwise noted; ruble pricing from CIS vendors is converted to 1 million tokens, and the dollar is shown for comparability at 80 RUB/$. For open-weight models, the API price from providers is shown as a reference or omitted: the key points are the license and the ability to self-host, and the economics are calculated through hardware TCO. “On-prem / CIS perimeter” means whether the model can be deployed inside the customer's perimeter.
Prices from non-Anthropic vendors are based on their public pricing; verify against the original source as of the relevant date (see the “Sources” section). Leading foreign models (Fable 5.1, GPT-6 Astra, Opus 4.8) are used under GDPR via KT.Team productized privacy gateway: personal data are anonymized before being sent to the cloud and restored in the response. The table shows price per token, but that is not yet the cost of the task.
| Model | Vendor | In/out price (1M) | Context | License | On-prem / CIS-based environment | Data in CIS / GDPR | Best use cases |
| Fable 5.1 | Anthropic | $10 / $50; cache reads $0.25 | 1M | Closed | No | Through a security gate (API proxy) | Long-lived agents: the same pricing, but cache reads cost four times less |
| GPT-6 Astra | OpenAI | $10 / $50; cache $1 | 1.05M (128K output) | Closed | No | Through a security gate (API proxy) | Interface work; access is rolled out in stages—check before planning |
| Fable 5 | Anthropic | $10 / $50 | 1M | Closed | No | Through a security gate (API proxy) | Only after checking availability: heavy long-horizon reasoning and agentic workflows |
| Claude Opus 4.8 | Anthropic | $5 / $25 | 1M | Closed | No | Through a security gate (API proxy) | Best price-to-intelligence default among leading closed models |
| GPT-5.5 | OpenAI | ~$5 / $30\* | ~1M+ | Closed | No | Through a security gate (API proxy) | Large context, cheap cache, and batch |
| DeepSeek V4 | DeepSeek | Flash $0,14 / $0,28; Pro $0,44 / $0,87 | 1M | MIT | Yes | Yes, if deployed in CIS | Code and long context within the customer's perimeter |
| Qwen 3.x | Alibaba | open-weight (Apache 2.0) | 128–256K | Apache 2.0 (junior); Max — closed | Yes (235B/Coder); Max — no | Yes, if deployed | Code, multilingual support, low-cost on-prem |
| GigaChat 3.5 Ultra | Sber | No API price; self-host/TCO | long context; 432B MoE | MIT | Yes | Yes, if deployed in CIS | A CIS open-weight model for on-prem, code, math, and agentic scenarios |
| Gemma 4 | Google | self-host / ~$0.06–0.30 hosted | 256K | open weights\* | Yes | Yes, if deployed | Low-cost high-volume inference in the perimeter |
| Llama 4 | Meta | self-host | 1M–10M | Community License\*\* | Yes | Yes, if deployed | Mature ecosystem, very long context |
| GigaChat API Max | Sber | 650 RUB / 1M ($8.1) | 128K | Closed/API | Cloud in CIS | Yes, data centers in CIS | CIS tasks without VPN and without your own hardware; ruble opex |
| YandexGPT | Yandex | ~200–$5 / 1M ($2.5–5.0) | 32K (Lite) / up to 128K (Pro) | Closed; 5 Lite — open (custom) | CIS cloud; Lite 8B — yes | Yes, claimed under Federal Law 152 | CIS tasks without VPN, payment in rubles |
* GPT-5.5 prices and context are based on public OpenAI statements; check the current OpenAI pricing page on the date in question. State specific multipliers (long-context threshold, regional markup) only with a link to the pricing page. * The Gemma 4 license allows commercial use, but historically it has not been fully OSI-open (there are use-policy restrictions). Before on-prem deployment, read the license text on HuggingFace. ** Llama 4 Community License is open-weight with restrictions (AUP, 700 million MAU threshold).
This is "open-weight with licensing restrictions," not classic open source.