72 hours: how a three-person team used Claude to breach OpenAI
Hacktron AI chained a libheif bug and an SSO flaw to reach OpenAI's internal codebase, with Claude Opus 5 building the exploit. The cost of a sophisticated attack is now set by model releases, not by hiring.

A three-person security team used Claude to break into OpenAI, and the whole thing took less than 72 hours from first bug to a pull request sitting inside OpenAI's internal codebase. The researchers at Hacktron AI disclosed it this week. OpenAI paid them a $6,500 bounty and patched the hole. The number that should stop you isn't the bounty. It's the 72 hours.
Here's what actually happened, because the chain is the story. The team found a heap buffer overflow in libheif, an image-decoding library that sits underneath an enormous amount of software. OpenAI's community forum runs on Discourse, which passed HEIC image uploads to ImageMagick, which called the vulnerable libheif. That gave the researchers a foothold. An SSO misconfiguration in OpenAI's identity layer then turned a forum compromise into access to ChatGPT and Codex accounts — and one of those accounts had Codex wired into OpenAI's GitHub organisation.
The model did the hard part
The part I keep coming back to is how the exploit got built. The team pointed Claude Opus 4.8 at the Discourse Docker image and asked it to look for security issues in the libheif package. It found that specific security fixes hadn't been backported. Building a reliable exploit with modern memory protections turned on is the genuinely hard, expert-only work — and Opus 4.8 struggled with it across several sessions.
Then Opus 5 shipped. They started a fresh session, and it produced a working exploit for a local Mac in about three hours, then ported it to the environment Discourse actually runs. Placed in an autonomous loop against the team's own test instance, it achieved remote code execution on its own.
A few things worth sitting with:
- The full campaign behind this — targeting libheif across Slack, Meta, GitHub Enterprise and others — cost under $3,000 in tokens and took three people two months.
- Adapting the exploit to each new target took one or two days, often starting almost blind to the target's exact configuration.
- Of all the companies hit during the research, only Shopify detected the activity, even after thousands of malformed images crashed image processors.
The researchers put it plainly: software has long been protected by security through complexity. Turning a known bug into a working exploit used to demand rare skill, real time and knowledge of the target. AI is converting that scarce expertise into compute. The defensive assumption that "someone would need a well-resourced team and months" is the assumption that just broke.
For anyone responsible for a security posture, the practical read is narrow and urgent. Audit your image-processing pipelines, patch libheif and libde265, and sandbox untrusted decoding. But the wider read is the one that changes planning: the cost of a sophisticated attack is falling faster than most threat models assume, and it's falling on a curve set by model releases, not by hiring.
Source: Hacktron AI — Hacking OpenAI