Google DeepMind has unveiled Gemini 4 Argon, a new frontier model designed for complex, long-horizon professional tasks. The model's most distinctive feature is an industry-leading 1 million token output limit, up from the previous 64,000 tokens, which allows it to sustain deep reasoning and generate hundreds of thousands of tokens in a single trajectory. According to the announcement, this expanded capacity is meant to let the model tackle tough problems in one go rather than breaking them into smaller steps.
Argon is already being used internally at Google, where thousands of employees are applying it to specialized coding, research, and writing tasks. The company reports concrete results, including a 2.7x speedup in a memory-safe video decoder after Argon agents replaced SIMD code with safe Rust, and memory optimizations across Google's data centers that freed up over 300 TiB of memory. The model also sets new state-of-the-art scores on benchmarks like DeepSWE v1.1 for software engineering, the Vals Index for economic impact, and LVBench for long video understanding.
On the safety side, Google is taking a phased approach. Argon is initially rolling out to a set of trusted cyber defenders through the company's Fairwind Program, and Google is participating in the U.S. government's voluntary pre-release model access process. The model will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input tokens priced at 95% off the input price. Google says it will continue gathering feedback from early testers before making Argon broadly available to developers, enterprises, and consumers.