Google DeepMind unveiled Gemini 4 Argon, its first larger-than-Flash model since February, aimed at coding, enterprise knowledge work, and cybersecurity. The company claims it leads on 13 of 19 published benchmarks against GPT-6 Astra and Claude Opus 5.5, with a DeepSWE score of 77.9% versus 74.2% for Opus 5.5 and 74.1% for Astra. However, the source notes that some observers questioned the published numbers, citing possible preference-data benchmaxxing.
Argon's headline feature is a 1M-token output limit, achieved through a new API feature called Long Decode Continuation that pauses and resumes long responses across calls. Independent evaluators differ: Vals lists a 262K max output, while Artificial Analysis reached 1M using the new feature. Pricing is $4/$20 per 1M input/output tokens, with a 50% introductory discount to $2/$10 and a 95% discount on cached input.
Access is initially restricted to government users and trusted cyber defenders in the Fairwind Program, with broader availability promised later. Internal reports claim Argon agents freed more than 300 TiB of data-center memory and are migrating over 800K lines of C/C++ kernel code to Rust. On agentic benchmarks, it ranks #1 on AutomationBench-AA at 77.5% but trails on Terminal Bench 4, and while its hallucination rate is lower than Astra's (15% vs 51%), its accuracy is also lower (50% vs 63%).