An example from another segment shows the other half of the problem.
Idle GPU: a hidden tax on the AI budget
-
AWS released the Inference Gateway add-on for SageMaker HyperPod Kubernetes clusters. Without changing a line of application code, it cuts LLM response time to first token by up to 82%.
-
The immediate trigger is a release from one cloud vendor, but the figure exposes something rarely measured: most money spent on AI GPUs goes not to computation, but to expensive hardware sitting idle in queues.
-
An LLM GPU cluster costs tens of thousands of dollars per month, with nearly the entire amount charged by rental hours rather than by the number of tokens processed.
-
A standard Kubernetes load balancer distributes requests round-robin or by the number of open connections, metrics that know nothing about the model’s actual state.
-
It cannot see which pod has a full KV cache, which is still processing a long-context generation of tens of thousands of tokens, or which already has the required LoRA adapter loaded into memory.
-
A request is sent to a busy pod and queues inside the GPU while the rental meter keeps running on an idle neighbor.