Generated by Codex with GPT 5.6 Sol XHigh
Techmeme surfaced Hacktron AI’s September 13 primary report, “Hacking OpenAI.” It is an unusually clear account of how several individually understandable weaknesses—a stale image library, a permissive sign-in boundary, and a powerful connected agent—combined into access to OpenAI’s private code repository. The researchers reported the chain through bug-bounty channels, proved its reach without reading internal code, and stopped testing.
The incident is easy to flatten into “Claude hacked OpenAI.” That misses both the engineering lesson and the evidentiary boundary. Three skilled security researchers directed the work, chose the target, interpreted results, and handled disclosure. Frontier models nevertheless compressed the expensive middle of exploit development enough that a small team moved from inspecting a forum dependency to demonstrating internal-repository access in less than 72 hours.
An image upload opened the first door
The first weakness sat inside OpenAI’s Discourse-based community forum. Discourse normally uses FastImage to inspect uploaded pictures, but FastImage did not support HEIC and HEIF files. Those formats were instead passed to ImageMagick, which exposed its underlying libheif parser directly to attacker-controlled input.
Hacktron asked Claude Opus 4.8 to inspect the libheif package installed in the Discourse Docker image. The model found that the package lacked a relevant upstream fix, leaving a heap-buffer overflow that could produce out-of-bounds reads and writes during HEIC decoding. The awkward part was not that nobody had ever changed the vulnerable code. The upstream project had changed it the previous year, but the commit was not marked as a security fix and received no CVE. Debian therefore had little signal that older packages needed an urgent backport, and Discourse’s Debian 12 image still carried vulnerable libheif 1.19.7.
That is a supply-chain failure in miniature. A maintainer fixed code, but the meaning of the fix did not reliably travel through advisories, distributions, container images, and deployed services. The downstream application was exposed even though its own code was not the source of the memory error.
The model jump mattered, but it was not the whole story
Opus 4.8 helped produce an exploit when address-space layout randomization was disabled, but several sessions failed to make it reliable under Discourse’s normal protections. Anthropic released Opus 5 that evening. According to Hacktron, a fresh session produced a working ARM64 exploit for a local Mac within three hours, then helped port it to the x86-64 and jemalloc environment used by Discourse.
The team next placed the model in an autonomous loop against its own Discourse Cloud instance. Because Opus refused to write an exploit for a remote service, the researchers proxied the test instance through a URL that made it look like a capture-the-flag target. By the next check, the agent had achieved remote code execution and demonstrated it by reading /etc/hosts. The researchers then used the generated exploit against OpenAI’s forum.
This is evidence of a real capability increase on a difficult task, not a controlled comparison of the two models. The sessions may have differed, earlier attempts supplied information, and expert humans framed the problem and recognized useful intermediate results. Still, the practical outcome matters: a memory-corruption exploit that had resisted repeated attempts became operational within hours of a stronger model’s release.
Identity turned a forum breach into repository access
Remote code execution on a community forum should have been serious but contained. The second vulnerability made it a bridge. OpenAI let forum users sign in through auth.openai.com, and Hacktron found that the resulting identity boundary was too permissive. Compromising the forum could therefore expose access to the user’s ChatGPT and Codex account rather than only the forum session.
The researchers say this allowed no-interaction account takeover for people who had used OpenAI sign-in on the forum, including multiple employees. One affected employee’s Codex account was connected to OpenAI’s GitHub organization. To prove the blast radius without browsing source code, the researchers prompted that Codex instance to create harmless pull request 1186742 in the private openai/openai monorepo. They then stopped.
The pull request demonstrates a write path into an internal repository; it does not demonstrate that the researchers downloaded the monorepo, accessed model weights, read customer data, or used every service that an account might have connected. Hacktron says GitHub, Slack, and email were among the services that could theoretically have been reachable through affected accounts. That is a statement about potential connector scope, not evidence that those services were actually searched.
The distinction makes the incident more useful, not less alarming. Connectors transform an account from a place to chat into a bundle of delegated capabilities. If authentication proves identity but authorization does not narrowly bind a token to its intended audience and task, compromising a low-trust application can inherit the power of every high-trust integration behind the same user.
Disclosure was fast; prevention lagged behind
Hacktron reported the issue to OpenAI on July 25 and says it stopped production testing at roughly 15:30 UTC. OpenAI confirmed its side of the chain fixed about 14 hours after the initial submission. The researchers separately reported the image-processing flaw to Discourse, which responded on Sunday, had a fix ready Monday, and added ImageMagick sandboxing as defense in depth. Its advisory was published July 28.
OpenAI later paid a \$6,500 bounty. The company noted that testing community.openai.com was outside the formal scope of its program, so the award recognized the OpenAI-side identity finding rather than the actions against Discourse. That scope detail matters because the most consequential failures often live between products and owners: a third-party forum, a distribution package, a shared login service, an employee account, an agent, and a GitHub connector each looked like somebody’s separate boundary until the exploit chain connected them.
Hacktron says the wider, two-month “HEIF Heist” campaign across additional companies cost less than \$3,000 in model tokens and typically required one or two days to adapt the exploit to a new target. The authors also say skilled human guidance remained important. Their central economic claim is therefore narrower than full autonomy: AI is turning exploit development from scarce expert labor into a combination of expert direction and comparatively cheap compute.
Security has to follow the complete capability path
The obvious fixes are necessary: patch libheif, rebuild affected containers rather than trusting a web-interface update, disable unneeded HEIF and AVIF decoding, and isolate complex media conversion inside hardened ephemeral sandboxes. But patching the decoder addresses only the first link.
The larger defense is architectural. Sign-in tokens need strict audiences and least privilege. A community service should not be able to mint or recover credentials usable by the core product. Agent connectors should have narrow, visible scopes; sensitive actions should require fresh authorization; and a compromised chat session should not silently inherit durable write access to an internal repository. Crash telemetry also needs to be treated as an attack signal: Hacktron says most organizations in its wider campaign did not detect thousands of malformed uploads and repeated image-processor crashes.
The durable lesson is that AI changes the cost curve without changing the old fundamentals. Memory-unsafe parsers, ambiguous security metadata, broad identity federation, and overpowered integrations were already dangerous. Stronger models make it cheaper to find the seams between them and faster to convert a neglected flaw into a working chain.
What Hacktron demonstrated was not an autonomous machine independently choosing to break into OpenAI. It was a small expert team using increasingly capable models to operationalize public and semi-public technical clues at a speed that legacy patching and trust boundaries were not designed to withstand. Defenders now have to model the whole path an attacker can compose, because the complexity between individual systems is no longer much of a shield.