Security researchers used Claude to help them hack into OpenAI
First reported by The Verge ·
AI language models can now find and exploit zero-day vulnerabilities in third-party services, extending the attack surface for major tech companies.
Three independent security researchers from Hacktron successfully compromised OpenAI's internal systems by exploiting a vulnerability in Discourse, a third-party service used for OpenAI's community forums. They utilized Anthropic's Claude Opus 4.8 and 5 models to gain Remote Code Execution (RCE) on Discourse Cloud, subsequently accessing OpenAI's "Monorepo" GitHub repository. This repository allegedly holds OpenAI's proprietary algorithms. While they did not access the internal code directly, they demonstrated their access by submitting a pull request from an employee's Codex account. The exploit, part of their "HEIF Heist" project, took minimal time and cost to adapt to various companies, including Slack, Meta, and GitHub. The vulnerability has since been patched, and Hacktron received a $6,500 bug bounty from OpenAI.
This incident highlights a significant evolution in cybersecurity threats, where advanced AI models are not just tools for defense but potent instruments for offense. The rapid adaptation of the exploit across multiple platforms suggests that AI-assisted vulnerability discovery and exploitation could become a scalable and cost-effective method for malicious actors. This blurs the lines between human ingenuity and AI capability, making it harder to attribute attacks and necessitating a re-evaluation of security perimeters.
The reliance on third-party services like Discourse, now proven to be susceptible to AI-driven attacks, exposes a critical systemic risk for large organizations. Companies need to reassess their vendor risk management and consider the potential for AI-powered exploits targeting their software supply chain. Future security efforts will likely need to incorporate AI-driven threat detection and simulation to stay ahead of these rapidly evolving attack vectors.
AI-written summary. May contain errors.