Sources: as Google prepares to roll out its Gemini 4 Argon, staff are testing a new version internally named Carbon; one staffer says it "feels like Opus 5.5"
First reported by Businessinsider ·
Google's internal Gemini 4 Carbon model shows capabilities potentially matching Anthropic's Opus 5.5 for coding.
Google employees are internally testing a new version of its Gemini 4 AI model, codenamed 'Carbon,' which is reportedly an advancement over the upcoming 'Argon' release. This internal testing, shared via documents and screenshots, indicates Google is rapidly developing its Gemini 4 series, having announced the models in late September. One staffer described Carbon as feeling comparable to Anthropic's most advanced coding model, 'Opus 5.5,' though further testing is needed. Google has been testing a series of Gemini 4 models internally, including 'Argon,' 'Barium,' and 'Carbon,' with codenames often differing from public releases. For instance, 'Barium-B' is expected to be publicly known as 'Argon.' While 'Argon' was reported to achieve frontier performance on certain benchmarks, including legal and finance, and focus on cybersecurity, internal feedback suggested early versions lagged in some coding tasks compared to older Anthropic models. The release of Carbon as a distinct model or an update to Argon remains undetermined, as Google declined to comment on its internal developments. This internal development race highlights Google's intensified competition with rivals like Anthropic and OpenAI in the AI coding agent market.
Google's rapid development of Gemini 4 models, evidenced by internal testing of 'Carbon' as an improvement over 'Argon,' signals an aggressive push to close the gap with AI coding rivals. The company is deploying multiple internal codenames like Barium and Carbon to iterate on its Gemini 4 family, indicating a focused strategy to compete at the frontier of AI capabilities, particularly in agentic coding tasks where Anthropic and OpenAI currently lead. Google's efforts suggest a direct response to market pressure, aiming to regain a competitive edge with advanced model performance.
This accelerated internal development and testing cycle means users could see faster iterations of AI models with enhanced coding and complex task handling abilities. The benchmark comparisons to Anthropic's Opus 5.5, even if preliminary, indicate a potential uplift in performance for Google's offerings, affecting developers and businesses relying on AI for engineering work. Future releases will likely reflect this intensified competition, with user experience and model effectiveness becoming key differentiators.
AI-written summary. May contain errors.