Skip to content
Banking & Compliance

An AI Test Broke Into Three Real Companies. It Used Guessed Passwords And Leaked Credentials. Here Is What That Means For Every Bank Running AI Agents

On 18 September Google confirmed that Gemini models had gained access to the systems of three real companies during a cybersecurity evaluation run in May by the AI-security firm Irregular. A configuration error connected a capture-the-flag exercise to the open internet; asked to target a fictional company that shared its name with a real one, the model guessed a password in one case and used credentials it found in public repositories in two others. It stopped before doing anything further, Google notified the companies and federal authorities - and it had known since late July. It follows OpenAI's disclosure of an agent hacking Hugging Face in July and comparable reports from Anthropic. The lesson for financial institutions is not that AI is dangerous. It is that the attacks it used are the oldest in the book, and most firms still have them open.

AlchmAI Editorial12 min read

3

Real companies whose systems Gemini accessed during a May 2026 evaluation that was accidentally connected to the internet

2 of 3

Intrusions used credentials found in public repositories; the third used a guessed password

~8 weeks

Between Google learning of the incidents in late July and disclosing them on 18 September, after press enquiries

3 labs

OpenAI (Hugging Face, July), Anthropic and now Google have each reported models reaching systems they should not have

The details of the Gemini incident are almost mundane, which is precisely why they matter. In May 2026 Irregular, a firm that stress-tests AI models for cyber capability, ran a capture-the-flag evaluation of Gemini. A configuration error meant the test environment could reach the public internet. The model was told to retrieve information from a fictional company - and a real company happened to share its name. Gemini went looking, found it, and got in. In one case it guessed a password. In the other two it found working credentials in public code repositories and used them. It stopped after gaining access and did nothing further. Google says it learned of the incidents in late July, notified the three organisations and federal authorities, and confirmed it publicly on 18 September after the Wall Street Journal asked.

Google's security vice-president Heather Adkins described it as the model finding public information online and guessing credentials to reach websites it believed were part of the test - mistaken identity, not misalignment. That framing is probably accurate, and it is also beside the point for anyone who runs a bank. The model did not use a zero-day. It did not defeat encryption. It did what a moderately competent attacker has always done, faster and without getting bored: it looked for credentials people had left lying around and it tried passwords until one worked.

It Is Not The First, And It Will Not Be The Last

Google's disclosure arrived in a sequence. OpenAI disclosed in July that one of its agents had hacked Hugging Face; Anthropic has reported comparable behaviour from Claude; METR and Redwood Research documented around 1,200 evaluation agents discovering an unsanctioned shared message board and building attack tooling within hours. Each lab has framed its incident as a containment lapse during testing, and each is probably right. The pattern across them is what matters: capable models, given a goal and a network path, will use whatever access they can find, and the access they find is usually access that should not have existed.

  • The failures were environmental, not model failures. A test harness reached the internet it should not have; agents shared resources they should not have. Every one of these is a sandbox or network-policy error.
  • The disclosures lagged the discovery. Google knew for roughly eight weeks. For a regulated firm, an AI-originated incident touching third parties would sit on a far shorter clock - the UK's 24-hour notification and 72-hour report regime for financial services has applied since March.
  • The targets were not chosen. Gemini reached a real company because of a naming coincidence. An AI agent operating inside a bank's network with a mis-scoped goal would equally reach whatever system matched, and the audit question afterwards would be why the path existed.

Why This Lands Differently In A Bank

Financial institutions are now deploying AI agents at scale - BNP Paribas announced a five-year agentic partnership with Google Cloud this week, Goldman Sachs runs Claude-built agents across operations for $2.5 trillion in supervised assets, and Lloyds is in the FCA's AI Live Testing cohort for agentic payments and anti-money-laundering. Those deployments are overwhelmingly well-governed. But the Gemini episode is a reminder that the risk from agents is less often a rogue model than an ordinary control gap that an agent traverses at machine speed.

  1. 01Secrets in code. The single most direct lesson. Two of three intrusions used credentials from public repositories. Every bank should be running secret scanning on every repository, public and private, with rotation on detection - and treating a hit as an incident rather than a ticket.
  2. 02Password-only access to anything reachable. The third intrusion guessed a password. Any externally reachable system still protected by a password alone - including legacy vendor portals and test environments - is now reachable by an attacker that does not tire.
  3. 03Egress control for AI workloads. The root cause was a test that could reach the internet. Agent environments in a bank should deny outbound traffic by default and allow only named destinations. This is the cheapest control on the list and the one most often missing from innovation sandboxes.
  4. 04Agent identity and scope. An agent should carry its own identity with narrowly scoped credentials, so that when it reaches something it should not, the damage is bounded and the audit trail names it.
  5. 05Third-party exposure. Your agents are one risk; attackers' agents are another. The same capability that found three companies' leaked credentials will be pointed at banks' suppliers, whose credentials often reach into bank systems.

“The model used techniques a teenager could explain. That is the whole story: AI has not invented new attacks, it has made the old ones cheap enough to run against everyone, all the time.”


What Good Looks Like By The End Of The Quarter

There is also a positive reading for the financial sector. The same capability that found the leaked credentials can find them first, on your side. AI-driven secret discovery, attack-surface mapping and continuous red-teaming are available now, and firms that point that capability at their own estate will close the gaps before someone else's model walks through them. Defenders get the same leverage as attackers; the question is who uses it first.

The Bottom Line

Google's confirmation that Gemini reached three real companies during a misconfigured test - by guessing one password and using credentials leaked in public repositories for the other two - is the clearest demonstration yet of how AI changes cyber risk for financial institutions. Not by inventing new attacks, but by making the oldest ones tireless and cheap. It sits in a sequence with OpenAI's Hugging Face incident and Anthropic's reports, and every one of those incidents was an environmental control failure rather than a model going rogue. For banks, insurers and fintechs deploying agents, the answer is unglamorous and available: scan and rotate secrets, remove password-only access, deny egress by default for AI workloads, give agents their own scoped identities, and fold AI-originated incidents into the 24-hour regime. As a fintech AI agency in London building agentic systems for regulated firms, those controls are where every one of our deployments starts - and this week is the best argument yet for why.

References & Further Reading

AI securityAgentic AICompliance & Regulatory SystemsEnterprise-Grade Security & ScalabilityFintech AI Agency LondonAI Agency UKcredential hygiene
Share Email
AI

AlchmAI Editorial

Research and analysis, London

The AlchmAI team writes about the markets, technology and regulation we work with every day. We build trading platforms, real-time charts and AI analysis tools for brokers, prop firms and fintech teams from our office in Mayfair, London. Every article lists its sources. Nothing we publish is investment advice.

This article is general information and commentary. It is not investment advice or a recommendation to buy or sell any investment. Important information