LLM capabilities in 2026: choosing the right model for your process and budget

Comparison of frontier, open-weight, and CIS LLMs in 2026 by inference cost, context window, licensing, on-prem, and Federal Law GDPR, including GigaChat 3.5 Ultra.

  • Not one best model, but a model for the process
  • Three model classes - choose by need
  • Comparison of 10 LLMs in 2026: price, context, license, on-prem
  • Important version notes (fact check)

June 27, 2026. Updated July 6, 2026: added the GigaChat 3.5 Ultra release.

Updated September 22, 2026: pricing for Claude Opus 5.5, DeepSeek V4, GigaChat API, and Alice AI LLM (Yandex), with the Central Bank exchange rate of $1/USD. “Which LLM is best?” is the wrong question. The right question is which model handles a specific process at the required quality, at the lowest cost per outcome, without violating personal-data requirements. In CIS enterprise environments, this is most often what stops a pilot: the model is selected but blocked from production because information security and legal teams have not approved transferring personal data.

Below is a comparison of ten current LLMs as of September 2026 by inference cost, context, licensing, on-prem deployment, and GDPR suitability, plus the methodology KT.Team uses to select a model for a process, calculate cost per outcome, and build a perimeter that keeps personal data out of foreign clouds.

Prices and exchange rate preserved as checked 22.09.2026; facts about GigaChat 3.5 Ultra were added based on the release 06.07.2026.

LLM data becomes outdated within weeks, so check the price against the primary source before calculating.

10 LLMcompared by price, context, license, and suitability for the CIS environment
3 deployment optionsforeign API, privacy gateway, or on-prem / CIS cloud
06.07.2026GigaChat 3.5 release added; prices must be checked before the pilot

Not one best model, but a model for the process

The LLM market in 2026 is not one leader, but a set of tools for different tasks.

Leading closed models (Fable 5, Claude Opus 5.5, GPT-5.5) deliver maximum intelligence for complex reasoning and long agentic tasks, but are expensive per token and unavailable for deployment in your own perimeter. Open-weight models (DeepSeek V4, Qwen, Gemma, Llama, GigaChat 3.5 Ultra) can be deployed on-prem with full data control, but their economics depend on hardware and GPU utilization rather than public API pricing.

CIS cloud LLMs (GigaChat API, Alice AI LLM) provide native GDPR compliance and accept payment in rubles.

Model selection is matching the process profile (task type, volume, data sensitivity, latency) with the model profile.

That is why this article is structured not as a ranking, but as a comparison table plus selection rules - let us start with three model classes.

Three model classes - choose by need

Frontier closed

Opus 5.5, GPT-5.5 (Fable 5 after access verification): maximum reasoning capability and long agentic tasks; on-prem deployment is unavailable, so use only through a privacy gateway.

Open-weight on-prem

DeepSeek V4, Qwen3-Coder, GigaChat 3.5 Ultra, Gemma 4, Llama 4: code and data inside the customer's perimeter, full control, economics through self-host/TCO.

CIS APIs

GigaChat API, Alice AI LLM: native GDPR compliance, processing in CIS data centers, payment in rubles, no VPN required. This starts faster but is not equivalent to self-hosting.

Comparison of 10 LLMs in 2026: price, context, license, on-prem

Prices per 1 million tokens (input / output), unless stated otherwise; CIS vendor pricing is converted to 1 million tokens, and the dollar comparison uses $1/USD (Central Bank of CIS rate on September 22, 2026). CIS cloud prices include VAT; dollar prices exclude VAT. For open-weight models, provider API pricing is indicative or unavailable: the key factors are licensing and self-hosting capability, while economics are calculated through hardware TCO. “On-prem / CIS perimeter” means whether the model can be deployed within the customer's perimeter.

Pricing from non-Anthropic vendors is based on their public price lists; verify it against the original source as of the relevant date (see the “Sources” section). Leading foreign models (Fable 5.1, GPT-6 Astra, Opus 5.5) are used under GDPR through KT.Team productized privacy gateway: personal data are anonymized before being sent to the cloud and restored in the response. The table shows price per token, but that is not yet the cost of the task.

ModelVendorIn/out price (1M)ContextLicenseOn-prem / CIS-based environmentData in CIS / GDPRBest use cases
Fable 5.1Anthropic$10 / $50; cache reads $0.251MClosedNoThrough a security gate (API proxy)Long-lived agents: the same pricing, but cache reads cost four times less
GPT-6 AstraOpenAI$10 / $50; cache $11.05M (128K output)ClosedNoThrough a security gate (API proxy)Interface work; access is rolled out in stages—check before planning
Fable 5Anthropic$10 / $501MClosedNoThrough a security gate (API proxy)Only after checking availability: heavy long-horizon reasoning and agentic workflows
Claude Opus 5.5Anthropic$4 / $20; cached input $0.20Check with the vendorClosedNoThrough a security gate (API proxy)Best default price-to-intelligence choice among leading closed models: according to the vendor, performs at Fable 5.1 level on most tasks
GPT-5.5OpenAI~$5 / $30\*~1M+ClosedNoThrough a security gate (API proxy)Large context, cheap cache, and batch
DeepSeek V4DeepSeekFlash $0.30 / $1.20; Pro $1.32 / $3.96 (peak hours; twice as cheap off-peak)1MMITYesYes, if deployed in CISCode and long context within the customer's perimeter
Qwen 3.xAlibabaopen-weight (Apache 2.0)128–256KApache 2.0 (junior); Max — closedYes (235B/Coder); Max — noYes, if deployedCode, multilingual support, low-cost on-prem
GigaChat 3.5 UltraSberNo API price; self-host/TCOlong context; 432B MoEMITYesYes, if deployed in CISA CIS open-weight model for on-prem, code, math, and agentic scenarios
Gemma 4Googleself-host / ~$0.06–0.30 hosted256Kopen weights\*YesYes, if deployedLow-cost high-volume inference in the perimeter
Llama 4Metaself-host1M–10MCommunity License\*\*YesYes, if deployedMature ecosystem, very long context
GigaChat APISberLite $1, Pro $6, Max $8 per 1M, input = output, including VAT ($0.8 / $6.0 / $7.7); minimum $7 per month128KClosed/APICloud in CISYes, data centers in CISCIS tasks without VPN and without your own hardware; ruble opex
Alice AI LLMYandex$6 / 1,200 per 1M including VAT ($6.0 / $14.3); YandexGPT Lite — $2 / 200Check with the vendorClosed/API; YandexGPT 5 Lite — open (custom license)CIS cloud; Lite 8B — yesYes, claimed under Federal Law 152CIS-language tasks without a VPN, payment in rubles; Yandex AI Studio flagship

* GPT-5.5 prices and context are based on public OpenAI statements; check the current OpenAI pricing page on the date in question. State specific multipliers (long-context threshold, regional markup) only with a link to the pricing page. * The Gemma 4 license allows commercial use, but historically it has not been fully OSI-open (there are use-policy restrictions). Before on-prem deployment, read the license text on HuggingFace. ** Llama 4 Community License is open-weight with restrictions (AUP, 700 million MAU threshold).

This is "open-weight with licensing restrictions," not classic open source.

Important version notes (fact check)

ModelClaimedReality
Qwen Maxopen-weightQwen3-Max and 3.7-Max are proprietary API-only Alibaba models; weights not released. Alibaba's open-weight options — junior Qwen3: 235B-A22B and Coder-480B (Apache 2.0); deploy these on-premises.
Llama (latest)open-weightLatest open-weight — Llama 4. Meta's newest model (Muse Spark, April 2026) — proprietary, no open weights; for open-weight comparison, Llama 4 (Scout/Maverick) is the correct choice.
Fable 5availableAnthropic, model ID claude-fable-5; the 2026-06-09 announcement is confirmed, but the 2026-06-12 update reports restricted access. It should not be added to a production shortlist without checking the current status with the vendor.
DeepSeek V4open weightsMIT license — the cleanest of all open-weight models in the table: deploys in CIS infrastructure without MAU restrictions.
GigaChat 3.5 Ultranew paid APIAs of 06.07.2026, these are published 432B weights under MIT, not a new GigaChat API tariff. We do not insert an API price into the calculator; for on-prem, we calculate hardware, GPU utilization, and DevOps.
GigaChat on-premon-prem applianceConfirmed: GigaChat API cloud processing in CIS data centers under GDPR and open-weight GigaChat 3/3.5 (MIT). Check with Sber before making public claims about delivering proprietary cloud GigaChat as an on-prem boxed solution.

Benchmarks - with caveats

Public benchmarks are good for shortlisting, but not for selection: they are vendor-picked, do not reflect your task, your prompt, or your acceptance criteria, and any “best model” table becomes outdated within weeks. You should rely on running candidates on your own tasks. What specific models claim:

Fable 5

Anthropic claims leadership on SWE-bench and financial benchmarks. The percentages come from the vendor; verify them in the official announcement and do not present them as independent facts. Separately verify access status on the pilot date.

DeepSeek V4

Claimed to be the strongest open-weight model for code. SWE-bench / LiveCodeBench / GPQA are in the model card and API docs; if V4 has already been released, the figures are preliminary.

GigaChat 3.5 Ultra

Sber claims improvements in code, math, and agentic scenarios relative to 3.1 Ultra, a smaller KV-cache, higher throughput under load, and faster greedy decoding thanks to MTP. These are vendor claims - use them for shortlisting, then verify on your own eval.

CIS language

GigaChat remains an important candidate for CIS-language tasks and the CIS perimeter: the API enables a fast start in CIS data centers, while GigaChat 3.5 Ultra adds an open-weight route for self-host.

Calculate result cost, not price per token

  1. Price per 1 million tokens is unit cost, not task cost.

  2. A model that is 5x more expensive per token can be cheaper per result if it solves the task in one pass instead of three and does not require manual cleanup.

  3. You need to compare the cost of a completed task at the required quality - including retries, manual cleanup, and the risk of failing acceptance.

  4. That is why we apply a rework coefficient to weaker models - the estimated share of tasks they do not pass on the first try relative to frontier models (roughly based on reasoning and code benchmarks).

  5. For frontier-class models it is roughly zero; for simpler models it adds a premium to the output cost, and we show that premium as a separate line instead of hiding it in the average price.

Price per token

  • simple unit metric from the price list
  • does not account for retries and manual cleanup
  • does not show the risk of acceptance failure

Result cost

  • T_in / T_out, iterations and success rate
  • batch, prompt caching, latency p95
  • Integration TCO, eval, and on-prem hardware

Preliminary estimate

Calculator: cost and payback of AI adoption

The main parameters are visible immediately. The model is deterministic: every number comes from a visible mechanism.

Company size

A preset seeds the starting values; every field stays editable.

Back office

How many roles and processes there are, and what they cost today.

Workload

How many tokens a single process consumes per month.

Your estimate

Order of magnitude for your parameters

Transparent model

How this is calculated

All figures are order-of-magnitude estimates, not a commercial offer.

  • Inference is modeled at the Claude Opus 5.5 price as a frontier-model benchmark, using the same token volume for every deployment contour. Prices come from the 2026 LLM comparison as of 2026-09-22; calculator amounts are displayed in USD at a fixed rate of 1 USD = 84 RUB.
  • Implementation cost is modeled only per process: 300,000 RUB per process at a 60% automatable share. If the share is lower or higher, process cost scales proportionally. The calculator has no one-off launch fee. The minimum development contract is 1.5M RUB: a programme of stages or a dedicated outstaff team.
  • Processes with high variability or strict legal responsibility are automated partially. That is normal.

Two pricing models

Two commercial models: fixed price for a working process, or an outstaff contract for a dedicated ai-native team. The inference contour (cloud, on-premise, or your perimeter) is selected separately. More detail: pricing approach.

Guide for 6 API models

The same parameters at the default prices from the comparison table, without caching: `base = 1M·P_in + 0.5M·P_out`. Ruble amounts use an exchange rate of $1/USD. Rework* is the estimated share of tasks the model does not complete successfully on the first attempt relative to frontier models (rough benchmark-based estimate; approximately 0 for frontier models). The “with rework” column = base × (1 + coefficient): even a cheap model with a high rework rate can lose on cost per outcome.

GigaChat 3.5 Ultra is not included here: no published API pricing is available, so it is calculated within the on-prem/TCO perimeter rather than as API dollars per day. We did not estimate the rework share for Alice AI LLM; compare it on your own tasks. The figures are estimates based on this table as a single source.

Model$/day base (≈RUB)Rework*$/day with reworkCapability note
Fable 5.1~$35,00 (~$35)0%~$19–35the same pricing as Fable 5; for repeated context, the cache cuts the bill
GPT-6 Astra~$35,00 (~$35)0%~$35,00interface work; access is rolled out in stages, so check before planning
Fable 5~$35,00 (~$35)0%~$35,00maximum long-horizon reasoning; use only after checking availability
GPT-5.5~$20,00 (~$20)0%~$20,00heavy reasoning and large context; use selectively
Claude Opus 5.5~$14,00 (~$14)0%~$14,00Leading class, cheaper than Opus 5 and Opus 4.8; according to the vendor, at Fable 5.1 level
Alice AI LLM~$13,10 (~$13)Not evaluated—Yandex CIS cloud, ruble pricing including VAT
GigaChat API Max~$11,61 (~$12)20%~$13,93CIS cloud, ruble pricing including VAT; weaker on complex tasks, so rework is included
DeepSeek V4 Pro~$3,30 (~$3)10%~$3,63Peak pricing; twice as cheap off-peak
DeepSeek V4 Flash~$0,90 (~$1)25%~$1,13cheapest candidate; on complex tasks, it more often requires rework

September 2026: It is the agent's memory, not the token, that got cheaper

  1. Two models launched during the first week of September, and for agent workflows, the important difference is not the one usually discussed. Claude Fable 5.1 (September 1) kept its predecessor's pricing: $10 and $50 per million input and output tokens.

  2. Something else changed: reading from the prompt cache fell from $1.00 to $0.25 per million, a fourfold reduction.

  3. According to the vendor's estimate based on real August workloads, this reduces a typical bill by about a quarter and highly agentic workflows by almost half. GPT-6 Astra (September 3) costs the same $10 and $50, provides a 1.05-million-token context with output of up to 128,000 tokens, and has significantly improved interface work: 72.6% on the offline subset of OSWorld 2.0 versus 65.7% for the previous version.

  4. An important practical caveat: this is OpenAI's first model with a Critical cybersecurity rating, so access is being rolled out in stages, starting with enterprise customers.

  5. Plan a process around it only after access has been confirmed for your contract. Claude Opus 5.5 (September 22) also reduced the token price itself: $4 and $20 per million input and output tokens versus $5 and $25 for Opus 5 (July 24) and Opus 4.8; cached input is $0.20.

  6. According to the vendor, the model performs at Fable 5.1 level on most tasks and returns responses 30% faster than Opus

  7. For processes where the choice was previously between Opus pricing and Fable quality, this is the first candidate to recalculate.

  8. Why the cache matters more than the per-token price specifically for agents.

  9. In a long-running process, an agent rereads the same material: procedures, the data schema, dialogue history, and tool descriptions. In a one-off request, this accounts for a few percent of the bill; in a long-running agent workflow, it makes up most of it.

  10. Therefore, when choosing a model for an agent, compare not the “price per 1M tokens” line, but the cost of rereading the context.

  11. That is exactly why two models with the same pricing produce different bills at the end of the month.

Assess where AI can deliver impact in your process

Calculator: $/day including rework

Choose a model - prices will be filled in from the table above, and the rework coefficient will add an extra line item. For frontier models it is about 0; for simpler models (GigaChat API Max, DeepSeek) it is greater than zero. We do not insert GigaChat 3.5 Ultra as an API preset without a published tariff: for it, calculate self-host/TCO through the on-prem setup. Any field can be adjusted for your own process: input and output in million tokens/day, prices per 1M, exchange rate, cache share, and rework percentage.

—
$ / day base
—
$ / day including rework
—
$ / month · 22 days (with rework)
—
$ / day without cache
input — · cache — · output — · rework —

Rework coefficient is the estimated share of tasks the model does not complete on the first try relative to frontier models (roughly based on benchmarks; for frontier-class models it is about 0). Rework cost = base × coefficient and is built into the separate "$ / day incl. rework" line. Cache reads are priced at 0.1× input price (a typical prompt caching multiplier). Batch (−50%), long context, and infrastructure TCO are separate multipliers and are not included in the calculation. Order-of-magnitude estimate, not an offer.

Multipliers that change the picture dramatically

-50%

Batch API

50% off token cost from Anthropic and most vendors. For overnight, non-latency-sensitive pipelines, that is a direct 2x savings.

0,1×

Prompt caching

Re-reading a stable prefix costs about 0.1x the base input price. Any changing byte - datetime.now(), unsorted JSON, a variable tool set - breaks the cache.

4-5×

Output is more expensive than input

Output is usually 4-5 times more expensive than input. A verbose model with long preambles is more expensive than a concise one at the same quality - the output format needs to be trimmed.

reasoning

Reasoning / thinking

Multiplies output tokens. On simple tasks, that is pure overspend: tune the effort to the task instead of setting the maximum by default.

long-context

Long context

Surcharges for long contexts can change the calculation. Fable 5.1 lists a 1M context without a premium; for Opus 5.5 and other vendors, check the threshold and multiplier in the pricing valid on the relevant date.

Model selection procedure by process

  1. 01

    Define the task

    Describe the acceptance rubric: what done means should be verifiable, not just look good.

  2. 02

    Run the candidates

    Take 2-4 models and one representative task set.

  3. 03

    Measure

    T_in, T_out, N_iterations, Success_rate, and latency p50/p95.

  4. 04

    Estimate production cost

    Account for batch, cache, TCO, and environment constraints where applicable.

  5. 05

    Select

    Compare the cost of the result, not the benchmark rank.

Personal data, on-prem, and GDPR

Not all processes and not all data fall under GDPR: the law governs personal data - anything that identifies a specific person (full name, phone number, email, passport, tax ID, SNILS, card number, IP address combined with other fields). Anonymized, aggregated, synthetic, and purely corporate data (product codes, technical logs, public texts) do not fall under it.

So the first step is not “you cannot use a foreign model”

, but rather “which exact fields in this process are personal data.” If the process includes personal data, sending it to a foreign LLM API is a cross-border transfer of personal data: under GDPR it requires separate legal grounds, and since 2025 turnover-based liability has been introduced, making this a board-level risk, not a finance-team fine. This is where most AI pilots in CIS fail: the model is chosen, but it never reaches production because legal and security teams did not approve the transfer of personal data.

Next are three mutually exclusive deployment setups; the cost of each is calculated in the calculator above and is not repeated here.

Foreign frontier via a privacy gateway

When to adopt

  • maximum reasoning with minimal capex and a fast start
  • the alignment table does not leave the CIS environment
  • provable to security and legal - anonymization before the cloud

When not to use it

  • cross-border transfer remains in anonymized form
  • measured detector recall and sign-off from security/legal
  • is not suitable where leaving the perimeter is itself prohibited by regulation

On-prem open-weight · CIS cloud

On-prem open-weight (DeepSeek V4, Qwen3-235B, GigaChat 3.5 Ultra, Gemma 4, Llama 4)

  • data stays within the perimeter, strict data residency is built in natively
  • GigaChat 3.5 Ultra provides a CIS open-weight option under MIT, but requires self-host infrastructure
  • anonymization is done inside the environment, and personal data never leaves the perimeter
  • we choose the model based on the hardware and load (1-2 nodes with modern GPUs), not the biggest one

CIS cloud (GigaChat API / Alice AI)

  • GDPR is handled natively without your own hardware - processing in CIS data centers, opex model, rubles, no VPN
  • Reasoning is below frontier level on complex tasks, tied to a CIS vendor
  • the 650 RUB per 1M line refers to GigaChat API Max, not self-host GigaChat 3.5 Ultra
  • on-prem is justified only at high GPU utilization (we calculate the number in calculator, not made up)

On-prem is not chosen to save on tokens

Pattern 1. Privacy Gateway

Personal data does not leave for the LLM API in raw form

Detect

Find personal datafull name, phone, email, tax ID, social insurance number

Classify

Classifyentity type and risk

Pseudonymize

ReplaceNAME_1, PHONE_2

LLM API

Processde-identified text only

Re-hydrate

Restore valuesalignment table in the CIS environment
Reversible pseudonymization: real full names, phone numbers, email addresses, tax IDs, SNILS numbers, passports, cards, accounts, and IPs are replaced with placeholders, the mapping table stays inside the customer's environment, and the proxy inserts the originals back into the response. Under the hood are Microsoft Presidio for detection and spaCy with custom NER for CIS personal-data formats. Pseudonymization is reversible, so such data remain personal data under the law: "anonymized" should not become a false sense of protection.

PII pipeline acceptance checklist

KT.Team service

LLM Gateway: models compliant with GDPR

Security gate / API proxy before a foreign LLM: pipeline detect → pseudonymize → re-hydrate (see the diagram above), the mapping table does not leave the CIS perimeter - a condition under which the customer's legal and security teams approve use of a frontier model.

  • Detect → pseudonymize → re-hydrate
  • Alignment table in the customer's environment
  • Provable to security and the regulator
LLM Gateway →

FAQ

FAQ

Which LLM is the best in 2026?

There is no single best model. For complex reasoning, the shortlist includes Opus 5.5 and GPT-5.5, with Fable 5 subject to availability on the pilot date; for coding among open-weight models: DeepSeek V4 / Qwen3-Coder / GigaChat 3.5 Ultra; for CIS-language tasks: GigaChat API or self-hosted GigaChat 3.5. The “best” choice depends on the process, volume, data sensitivity, and budget; compare cost per outcome on your own tasks.

Can personal data be processed through foreign LLMs?

Directly - no: this is a cross-border transfer of personal data under GDPR with turnover-based liability from 2025. Safe only through an anonymization layer (security gate / API-proxy) or by hosting the model within your environment. KT.Team delivers both options end to end with a pipeline that can be demonstrated to security and regulators (AI for business).

Which models can be deployed on-prem in a CIS perimeter?

Open-weight: DeepSeek V4 (MIT), Qwen3-235B/Coder (Apache 2.0), GigaChat 3.5 Ultra (MIT), Gemma 4, Llama 4 (Community License with restrictions), and previous open-weight GigaChat 3 models. Closed models (Fable 5, Opus 5.5, GPT-5.5) cannot be deployed on-prem.

How do you calculate the real cost of inference?

Not by token price, but by cost per result: cost of one attempt `(T_in × P_in + T_out × P_out) × N_iterations`, divided by `Success_rate`. This includes batch (-50%), prompt caching (~0.1x for cache reads), and hidden TCO. For GigaChat 3.5 Ultra, use self-host/TCO because the API price has not been published.

What should you choose for CIS-language tasks without VPN?

Without your own hardware: GigaChat API (Sber) or Alice AI LLM (Yandex): processing in CIS data centers, payment in rubles, and stated GDPR compliance. For strict on-prem deployment: open-weight GigaChat 3.5 Ultra (MIT), DeepSeek V4 (MIT), or YandexGPT 5 Lite (custom license; review the terms).

Why is open-weight not always cheaper than a cloud API?

The license is free, operation is not. Self-host pays off at consistently high GPU utilization; with uneven or low load, hardware idle time makes cloud API cheaper. TCO based on the actual load profile is what matters, not the license price.

Conclusion

There is no single best LLM - there is a model for a specific process. Frontier closed models provide peak intelligence, open-weight models provide data control and on-prem deployment, and CIS APIs enable a fast start under GDPR without your own hardware. GigaChat 3.5 Ultra strengthens the CIS open-weight stack, but it does not replace an API price in the calculator: its economics are self-host/TCO.

The winner is whoever can do three things: match the process to the model, calculate cost per result (not per token), and handle personal data correctly - anonymizing before the cloud (security gate) or by deploying it within your own environment.

Sources

Verified: September 22, 2026

Discuss the article: LLM capabilities 2026: what to choose for...

Enter your email or phone number so we can get back to you.

Send via: