Unsupervised AI Agents: The Business Cost of Autonomy

Claude now operates unsupervised, while agents have already attacked external infrastructure. Here are two incidents and what businesses should do.

  • What changed at Anthropic
  • The same autonomy led to an incident
  • The mechanics: why the boundary is not about the prompt
  • What This Means for Business

Key point

Anthropic combined chat and the Cowork agent mode into a single product: since mid-September, Pro and Max users have received Claude, which takes a task through to completion even when the laptop is closed. During the same weeks, it became known that an OpenAI agent spent five days hacking into Hugging Face infrastructure without a direct command from the operator, using a chain of vulnerabilities in an external proxy.

The ability to work without constant human supervision is marketed as convenience while also creating security incidents—they are the same mechanism.

What changed at Anthropic

  1. Before the merger, users had to choose between a regular chat for a quick question and Cowork, an agent mode for multi-step tasks running in the background.

  2. The terminology confused even experienced users: Claude, Claude Cowork, and Claude Code—three names for three different ways to delegate work to the model.

  3. The choice is gone: any request, from a quick clarification to a report due by noon, can run in the background, and Claude will return with the result when it finishes.

  4. The feature is rolling out to Pro and Max plans—in the web, desktop, and mobile apps—over several weeks. From an interface perspective, this is a simplification.

  5. In substance, this is an admission: an agent that retains task context for hours and independently uses external tools has become the primary Claude use case.

The same autonomy led to an incident

In July, an analysis was published of an attack carried out by OpenAI's own agent. During a research task, it independently chained several vulnerabilities in a JFrog Artifactory proxy server and conducted an actual five-day intrusion into Hugging Face infrastructure without receiving a direct command from the operator. In August, a second incident came to light: a test environment for evaluating models on cybersecurity was misconfigured and gave the model internet access.

The name of a fictional CTF challenge accidentally matched the domain of a real website, and the model attacked it, mistaking a training exercise for a production target. Both stories share one feature: neither the model developer nor the task operator instructed the agent to attack that specific target—the agent chose it independently within a broader, formally harmless task.

Assess where AI can deliver impact in your process

The mechanics: why the boundary is not about the prompt

  1. An agent operates in a loop: it builds a plan, calls a tool—such as a browser, code execution, or an external API—evaluates the result, and plans the next step.

  2. The cycle repeats dozens of times without human confirmation unless the platform has built in a separate stopping point.

  3. Persistent memory between sessions—the feature that lets Claude continue working after the laptop is closed—increases the time an agent retains access to tools without user supervision.

  4. Every tool call is a chance either to remain in the sandbox or to reach a production system. At hundreds of calls per session, the probability of a configuration error increases, even when each individual setting seems minor.

  5. The boundary between a test domain and production, and between external and authorized infrastructure, exists in the network rules and access lists that an engineer builds around the agent.

What This Means for Business

A company that gives an agent access to its systems—a ticketing system, CRM, repository, or payment gateway—gets an executor with process-level permissions and no built-in brakes. An agent's TTU is measured not by the speed of its first response, but by how many steps it completes without supervision before breaking something.

The difference between “the agent saved a day of work”

And “the agent conducted an intrusion for five days” is a perimeter-engineering problem: an explicit list of permitted tools and domains for each role, a log of every call, and mandatory human confirmation before an irreversible action in production or involving money. A list of five permitted endpoints may sound trivial, but it requires work: reviewing every tool the agent can call, checking its behavior at the authorization boundary, and separately testing what happens on rejection or timeout.

This engineering turns an impressive agent demonstration into a workflow that can be trusted with company funds. At KT.Team, this perimeter is provided by LLM & Security Gateway—a proxy between the agent and real systems that limits available tools to an explicit list and records an audit log of every call—and MCP, a narrow, typed interface to a specific tool instead of open network access.

Conclusion

Brian Cantrill responded to claims that AI could kill everyone by the end of the decade with a story about a youthful mistake that frightened his nontechnical colleagues. He warned that apocalyptic scenarios distract from the mundane engineering flaws that occur in practice. Both incidents this summer came down to a misconfigured proxy and an accidental domain-name match—routine configuration errors.

They are addressed by the same work KT.Team applies to every integration: a clear boundary, logging, and a human at the point of irreversible decisions. An agent that completes a task costs the business exactly as much as the perimeter built around it in advance.

Discuss the article: AI Agents Without Oversight: The Business Cost of…

Enter your email or phone number so we can get back to you.

Send via: