According to a WIRED report, three AI engineers at the startup Axiom used a general-purpose language model to navigate a 2024 Toyota Corolla through an In-N-Out drive-thru in the Bay Area. The model, OpenAI's GPT-6 Astra, was linked to windscreen-mounted cameras and the car's power steering, with a safety driver ready to brake. The engineers said the model drove slowly but steadily to the pickup window, needing no prior coaching for the task.
The stunt is part of a broader set of experiments suggesting that text-focused AI models are developing a rudimentary understanding of the physical world. The engineers also built a benchmark called DrivingBench, where only Astra completed a simple parking-lot course, and at a very slow pace. Other models—Claude Fable 5.1 and Grok—managed only 45% and 11% of the course, respectively.
The engineers believe the driving skill is an emergent capability from training on multimodal data like images, video, and 3D models, rather than from explicit driving instruction. They also caution that putting a general-purpose model in control of a two-ton vehicle is a high-stakes undertaking. The article additionally highlights a new benchmark, Humanity's Sixth Sense, developed by Elorian AI and Scale AI, to measure models' intuitive understanding of physical scenes—a capability seen as essential for home robotics and other real-world applications.