Infrastructure for AI Agents: The End of Old DevOps

Modal raised $355 million and reinvented AI infrastructure with GPU snapshots, multicloud, and seconds instead of minutes. Learn what it means for business.

  • Why Old DevOps Cannot Handle Agents
  • Agent experience—a different contract with hardware
  • How it works: GPU snapshots and multi-cloud
  • What This Means for Business

Trigger

Modal—an infrastructure company for running AI models—raised $355 million in a Series C round and reoriented its architecture around agent experience: a hardware contract for processes that live for seconds. The company's CTO, Akshat Bubna, describes the redesign concretely: elastic inference, GPU state snapshots, isolated sandboxes for every run, and operation across 17 cloud providers at once. An agent is a short-lived process: it starts, runs for seconds, and ends in thousands of copies per hour.

Why Old DevOps Cannot Handle Agents

  1. All modern cloud infrastructure is designed for humans: one long-lived session, predictable workloads, and minutes of waiting in CI or during deployment—an acceptable price.

  2. Agent workloads are different: thousands of short, isolated runs per hour, each requiring a clean environment; a GPU is often needed for just a few seconds and then released immediately.

  3. A conventional VM takes a minute to start, a container takes tens of seconds, and a cloud GPU queue adds just as much again. A company running an agent product on this infrastructure must choose between two losing options: keep a pool of resources warm but idle and overpay, or endure a cold start lasting several minutes for every agent call.

  4. In the second case, TTU—the time to outcome—is dominated by infrastructure wait time: the model's seconds disappear into minutes of cold-start delay.

Agent experience—a different contract with hardware

At Modal, agent experience means scaling to zero and back without delay, an isolated sandbox for every task, and GPU access within seconds. A load test puts this to the test: can the platform handle one million short agent calls per day without cost or response-time degradation?

Assess where AI can deliver impact in your process

How it works: GPU snapshots and multi-cloud

The core technique is GPU snapshotting: the state of an already-warm process, loaded model weights, and initialized CUDA context are frozen and restored in seconds instead of reinitializing the GPU and fully reloading the weights, which can take more than a minute.

This is the engineering behind the phrase “a simple, fast agent”

- from the outside, the call appears instant; inside, it relies on a complex state-recovery mechanism. Multi-cloud across 17 providers solves the second part of the problem: burst agent workloads get a GPU wherever one is available at that exact moment, without hitting a single cloud's quota. Sandboxes physically isolate every agent run—if an agent executes code or interacts with external systems, the side effects of one run do not carry over to the next.

What This Means for Business

  1. As models become smarter, the harness around them—call orchestration, memory, and action control—shifts the human role from line-by-line model supervision to attention management: deciding where human review is needed.

  2. The infrastructure and orchestration beneath this harness only become more important: the more autonomous the agent, the higher the cost of every second of downtime and the more expensive a poorly designed layer between the agent and real actions in company systems becomes.

  3. For a company evaluating an agent project, the right question is not “which model,” but what latency and cost profile will result under bursty, concurrent agent workloads—and who owns the infrastructure when it fails.

  4. The agent's brain is almost always externalized into the Claude, GPT, Gemini, YandexGPT, or GigaChat API. Connecting that brain to real application actions—through MCP servers, an LLM & Security Gateway, model routing, and secure access to production data—is weeks of engineering work that companies consistently underestimate at the outset.

  5. This is precisely where KT.Team delivers projects: at the intersection of agent integration with real systems, rather than model selection. AI-native integration on top of an existing 1C, Pimcore, or Bitrix system requires the same discipline as Modal's GPU snapshots—unnoticeable from the outside, demanding under the hood.

Conclusion

$355 million to make a GPU wake up in one second instead of one minute sounds modest until you multiply that second by millions of agent calls per day. A company that does not solve this engineering problem solves it by paying for more idle GPUs.

Discuss the article: AI Agent Infrastructure: The End…

Enter your email or phone number so we can get back to you.

Send via: