The breach began on July 25 through a memory vulnerability in libheif, a library used by the Discourse forum software to process image uploads. Because the specific security patch for this bug lacked a formal CVE designation, the software remained unpatched, allowing the researchers to craft a malicious image file that hijacked the server. Once inside the Discourse environment, the team chained this flaw with a secondary exploit to gain access to OpenAI employee credentials, including accounts linked to the company’s internal GitHub organization.
Security Researchers Used Claude to Breach OpenAI Infrastructure
A team of three security researchers from startup Hacktron AI successfully compromised OpenAI's internal systems by leveraging Anthropic’s Claude Opus 5 model. The attack, executed as part of an official bug-bounty program, exposed critical vulnerabilities in third-party software and granted the team unauthorized access to employee accounts.

What makes this incident particularly notable is the role of AI in the exploitation process. The researchers reported that an earlier version of Claude, Opus 4.8, failed to construct a working exploit. However, within hours of Anthropic releasing Opus 5, the model successfully generated the necessary code to execute the attack. OpenAI has since resolved the identified issues and awarded the team $6,500 for their findings. This event underscores a growing concern among security experts: the rapid evolution of large language models is significantly lowering the barrier to entry for cyberattacks, allowing those with limited technical expertise to identify and weaponize software vulnerabilities that previously required months of dedicated research.



Comments (0)
No comments yet. Be the first!