A Separate Challenge: Monitoring Agent Systems
AWS described a case where an agent lacks IAM permissions to call the model, the response is empty, but no 500 error occurs - from the infrastructure perspective, everything is green.
AWS examined a flaw in LLM cost estimates: the price per token is not the price of the result. We explore accuracy, prompt caching, and silent agent failures.
AWS recently dissected a question every company gets wrong: how to choose a model for production.
The standard approach is to open the price list and compare dollars per million tokens.
There is one figure on every website, so that same figure ends up in the procurement table.
Businesses pay for a resolved support ticket, a finished research report, or an accurate figure in a financial summary - tokens are only an intermediate unit here.
Between the price on a vendor's website and the result delivered to the client are three multipliers that the price list does not show: how often the model answers correctly, how many tokens it needs to reach the right answer, and - for agent scenarios - how many steps it takes.
Each agent step resends the entire expanded conversation context to the model.
A model that is 30% cheaper per token can cost 200% more per solved task if it needs twice as many attempts and three times as many steps.
AWS measures three parameters: first-attempt accuracy, the number of tokens per correct answer, and the number of steps in an agent scenario. The real cost formula is the price per token multiplied by the token volume required to reach the result and the number of iterations, adjusted for the share of answers that had to be reworked. A company that evaluates models using only the first of these three parameters regularly chooses the wrong model.
There is also a less obvious cost: latency and repeated computation. A typical LLM request has a fixed part (instructions, documents, and conversation history) and a variable part (the user's actual question). For a support bot, the instructions may take 3,000 tokens while the customer's question takes 50. Without cache reuse, the model recalculates those 3,000 tokens from scratch for every request.
AWS showed that prefix-aware routing in SageMaker Inference - directing requests with the same prefix to the same instances - reduces latency and computational load without changing the model itself. This is the same principle behind prompt caching in LLM & Security Gateway: the same system instruction should not be recalculated a thousand times a day.
AWS described a case where an agent lacks IAM permissions to call the model, the response is empty, but no 500 error occurs - from the infrastructure perspective, everything is green.
Another scenario: a supervisor agent with a poorly formulated prompt starts routing 20% of requests to the wrong specialized agent, while infrastructure metrics remain unchanged.
Traditional monitoring of failures and delays does not detect these issues - a separate assessment of the agent's effectiveness is needed, not just its availability.
For the business, this means that 20% of customer requests are handled incorrectly while the dashboard shows 100% uptime.
A manager afraid of wasting the budget usually compares vendor price lists instead of calculating the cost of a solved task using their own data. The right test is to run real company cases through several models and calculate the result price: the cost of obtaining a correct answer, including rework and extra steps. The price of one request does not show this amount.
This is exactly the engineering layer KT.Team builds for clients through LLM & Security Gateway: a single routing point between OpenAI, Claude, GigaChat, YandexGPT, and Qwen, with prompt caching and metrics for solved tasks. RAG reduces the number of tokens and steps: the model gets context from a knowledge base instead of making it up from scratch every time. Evaluating agents against business metrics rather than HTTP statuses catches the 20% of quietly incorrect answers that standard monitoring misses.
A seemingly simple model choice is an engineering task: the cost of the result must be calculated with adjustments for accuracy, token volume, and the number of steps. Companies that continue comparing models by price list pay for the illusion of savings - and receive the bill in another month as low accuracy and extra rework.