Anthropic says it found 5,500 verified vulnerabilities in April-October and Glasswing partners found 129K+ in April-July; 33K+ were critical or high severity
First reported by Reuters ·
AI models now have a verifiable track record of security flaws, impacting their readiness for widespread deployment.
Anthropic has revealed that its bug bounty program, which allows external security researchers to test its AI models with reduced safety restrictions, identified 5,500 verified vulnerabilities between April and October. Concurrently, its partner Glasswing discovered over 130,000 vulnerabilities in the same April-July period, with more than 33,000 classified as critical or high severity. These findings stem from Anthropic's 'attack surface' program, an initiative designed to proactively uncover weaknesses in its AI systems by simulating adversarial attacks.
The sheer volume of vulnerabilities found, particularly critical ones, highlights the inherent security challenges in developing and deploying large language models. Anthropic's move to a more open testing environment suggests a growing recognition within the AI industry that rigorous, external security auditing is no longer optional but a necessity for building trust and ensuring safety.
This approach signals a maturing cybersecurity posture for AI developers, moving beyond internal checks to embrace a crowdsourced model for vulnerability discovery. The implications extend to users and enterprises, who can now anticipate AI systems undergoing more intense scrutiny, potentially leading to more robust and secure AI products in the near future.
AI-written summary. May contain errors.