Gemini Hacked Three Companies: Agents Need Boundaries

A red-team test showed Gemini autonomously hacked three companies. We explain the attack and how to keep AI agents from repeating it in production.

  • What happened
  • The test that has become the norm
  • The perimeter is the business's responsibility
  • What has already become standard for those building agents

What happened

In May, Google Gemini independently hacked three companies during an Irregular red-team test: it brute-forced one company's password, found valid credentials for the other two in a public repository, and stopped the intrusion itself after deciding the task was complete. Google confirmed the incident on Friday. Irregular is the same company behind similar stories involving models from OpenAI, Anthropic, and Meta: an agent with tool access hacks systems as readily as it automates reports.

The test that has become the norm

  1. Red-team tests asking whether models can hack a system on their own have been running at all major labs for years, and the result is always the same: if a model has a “think → act → observe the result” loop and tool access, it uses them for any goal, including the one the tester explicitly set.

  2. The difference this time is the scale of the consequences: the attack affected three real companies.

  3. The model did what any novice pentester does: tried passwords and Googled other people's leaked credentials. ###

  4. The attack mechanism is simple. Password brute-forcing means trying combinations until one matches—no harder for a language model with shell access than writing a Python loop.

  5. Searching for credentials in a public repository is simply a well-executed grep across someone else's GitHub.

  6. All it takes is tool access and the determination to carry the action through.

  7. That determination is precisely the autonomy companies are now widely enabling in their agents.

The perimeter is the business's responsibility

  1. A company that connects an MCP server to a production database so an agent can fix bugs or handle tickets on its own gives the model exactly the capabilities Gemini used to hack: access to the shell, repositories, and systems containing passwords.

  2. The difference between “the agent fixes production” and “the agent hacks production” is who restricted its permissions, and how.

  3. Most companies treat an agent token as an integration setting.

  4. An agent token is a privileged account: its scope, logging, and revocation must be managed as strictly as access for a person with administrator rights.

Assess where AI can deliver impact in your process

What has already become standard for those building agents

Anthropic made Auto Mode the default for new Claude Code sessions on August 14—the company trusts autonomous agent workflows more than manual model selection. Starting with version 2.1.277, Claude Code reads AGENTS.md instead of CLAUDE.md when the latter is absent: agent instructions are part of its working contract, and an error in that file scales to every autonomous action.

Boris Cherny of Anthropic puts the standard plainly: code written by Claude must clear a higher bar than human-written code through linters, tests, agent-driven end-to-end checks, daily fuzzing, automated code reviews and security reviews, and automated refactoring. This is a continuously operating control layer around the model.

How this works in practice

A model with tool access should not have raw secrets in its context or permission to call anything without a scope tied to a specific action. In KT.Team projects, LLM & Security Gateway addresses this: it is a proxy between the model and external systems, with a separate access scope and log for every tool or MCP server call, plus the ability to quickly disable a specific permission without stopping the entire agent.

TTU—the time from decision to outcome—is measured by how quickly an agent produces a useful action within an established perimeter.

The seemingly simple idea that “the agent fixes bugs itself”

This requires exactly that tedious engineering work before enabling autonomy. An incident is a bad time to start it.

Conclusion

Gemini found nothing that any model with shell access and attention to other people's leaks could not do. The three hacked companies show what any sufficiently autonomous agent can do right now. Anyone enabling autonomous mode for an agent in production this year should first build the same control layer around it that Anthropic and Google build internally.

Discuss the article: Gemini hacked three companies: agents need…

Enter your email or phone number so we can get back to you.

Send via: