AWS has launched model caching for Amazon SageMaker HyperPod—a mechanism that removes the gap between “requested an inference pod” and “the model responds.” For large models such as DeepSeek-R1 (600+ GB of weights), this gap means half an hour of downtime: details in AWS blog.
The main point extends beyond a single release: AI outcomes depend not on the choice of model, but on the engineering around it—specifically, how much time passes from startup to the first useful answer. This is TTU, time to use, and it is this metric that businesses count in money, not tokens. Before the model answers the first request, two downloads take place behind the scenes: the inference server image is pulled from Amazon ECR, followed by the model weights from S3, FSx for Lustre, or HuggingFace Hub.
For a compact model, this takes a couple of minutes. For DeepSeek-R1 or a similarly sized Llama, it takes at least half an hour. During that entire time, not a single pod is physically ready to handle production load, while customer traffic either waits or falls back to a reserve node that also had to be warmed up in advance.


