Artificial Analysis says Gemini 4 Argon (high) matches GPT-6 Astra (max) on its Intelligence Index and has a 15% hallucination rate, compared with 51% for Astra
First reported by Artificialanalysis ·
The benchmark for AI hallucination rates is set substantially lower, changing the default reliability of high-tier AI models.
Google DeepMind's new Gemini 4 Argon model has achieved parity with OpenAI's GPT-6 Astra on the Artificial Analysis Intelligence Index, both scoring 53. Gemini 4 Argon (high reasoning) notably exhibits a significantly lower hallucination rate of 15% compared to GPT-6 Astra's 51%. At its current 50% promotional discount, Gemini 4 Argon costs $1.99 per Intelligence Index task, making it 60% of the cost of GPT-6 Astra, though this price is expected to rise after the promotion ends. The model also demonstrates improved agentic capabilities, ranking high on benchmarks like AutomationBench-AA. Gemini 4 Argon is currently in a limited rollout to select users, featuring a 1 million token context window and multimodal inputs with text output.
Google's return to the top tier of AI labs with Gemini 4 Argon signifies increased competition and potential for rapid advancement in model capabilities, particularly in reducing hallucinations. The improved agentic performance suggests a move towards more robust AI assistants capable of complex task execution. This development could pressure competitors to focus on similar improvements, especially in areas where their models historically lagged.
The competitive pricing, especially with introductory discounts, indicates a market strategy to gain share by offering comparable intelligence at a lower cost. As the discount period concludes, the true cost-effectiveness of Gemini 4 Argon will be a key factor for adoption. The model's advanced features like the 1M token context window and Long Decode Continuation offer new possibilities for complex, long-running AI tasks, potentially redefining what users expect from advanced AI services.
AI-written summary. May contain errors.