Autonomous AI Agent: 27 Minutes of Value and a Security Hole

An agent built a route on its own in 27 minutes—and the same autonomy became a vulnerability in Claude. How businesses can mitigate the risks of AI agents with data access.

  • The agent built the route itself — and decided where else to go
  • What actually happened here
  • The same capability is the source of the vulnerability
  • Companies give agents access faster than they assess the risks

The agent built the route itself — and decided where else to go

A user asked ChatGPT with GPT-6 Astra to build 5- and 10-kilometer loop running routes from home using OpenStreetMap data. The agent worked for 27 minutes: it found the address through Nominatim, downloaded a map of roads and trails through Overpass, calculated the loops, and delivered the result in three formats — a visualization, GPX for watches, and GeoJSON for maps. The developer did not participate: the agent built the call chain itself, from defining the task to producing the finished file.

What actually happened here

  1. 27 minutes: the time it took the agent to pass through several external services, each with its own response format and limitations, deciding what to do next at every step.

  2. Previously, the developer built this chain: writing code for Nominatim, working out the Overpass format, and manually calculating the loop geometry.

  3. Now the agent builds it on the fly for a specific person's specific request.

  4. For businesses, this means that time to usefulness (TTU) for tasks of this kind has fallen from weeks of development to a single request.

  5. The model learned to reliably call tools in sequence and verify its own intermediate results — this level of engineering maturity is now built directly into the product.

The same capability is the source of the vulnerability

  1. The mechanism that creates value — the agent deciding which external resources to call and in what order — turned out to be a security hole.

  2. A researcher made Claude's web_fetch tool follow a chain of nested links on a malicious site and exfiltrate user data. Anthropic fixed the vulnerability, but the attack principle applies to any agent with network access: if the agent decides which links to follow next, a page can override that decision.

  3. Separately, Anthropic ran the Claude Mythos model for 60 hours to search for weaknesses in cryptographic primitives — HAWK and weakened AES — with almost no human involvement.

  4. The findings do not currently represent a practical threat: they are a hypothesis about future risk, not confirmed evidence of an attack today.

  5. But the capability is the same: a model with autonomy and tools finds things that were not explicitly designed into it. In one case, the agent found a running route. In another, a path to someone else's data.

  6. The mechanism is the same: an agent with access to the outside world explores the space of possibilities faster and more broadly than a human, including areas the developer never intended to expose.

Assess where AI can deliver impact in your process

Companies give agents access faster than they assess the risks

Companies are widely connecting agents to email, CRM, internal APIs, and web search, inspired by demos such as the running route: look how quickly the agent works. The questions of what exactly the agent may read and call, and who verified it, often go unanswered. A developer who connects an MCP server to an internal customer database without an explicit list of permitted operations repeats the same design error that caused the web_fetch security hole — only now the company's data is at stake, not someone else's website.

A leak through an agent with access to CRM or email has a tangible cost: hours of investigation, legal risk, and lost customer trust. When no one limits the agent's autonomy in advance, there is no one to hold accountable after an incident.

How to address it through engineering

  1. Restricting agents' tools is a poor answer: it loses value such as the route example.

  2. The agent's capability boundary must be designed in advance as an explicit, verifiable structure. ### Allowlist of permitted calls

  3. The agent gets access to specific MCP tools and specific domains; everything else is blocked by default. ###

  4. The LLM & Security Gateway between the model and the outside world logs and filters every tool call before it goes out. This layer stops the agent from following someone else's link or making a call somewhere other than intended. ###

  5. A complete audit of the call chain from 27 minutes of agent operation — dozens of intermediate decisions.

  6. The company stores the entire chain, including intermediate steps; without it, investigating an incident after a data leak is impossible.

  7. This is part of the development platform — the control layer between "wrote a prompt" and "running in production" that makes agent autonomy manageable.

Conclusion

The agent that built a running route in 27 minutes and the agent that could be used to extract someone else's data belong to the same class of system. Agent autonomy is real in 2026 and is already delivering business value, but it requires an engineering control layer: an allowlist, a gateway, and a log of every step. A company that connects tools to an agent without this layer gets a running-route demo with unassigned risk included.

Discuss the article: Autonomous AI Agent: 27 Minutes of Value and…

Enter your email or phone number so we can get back to you.

Send via: