A group of Bay Area tech workers calling themselves DrivingBench connected several off-the-shelf large language models to a rented Toyota Corolla and let them drive around a cone course in a parking lot. The experiment, reported by 404 Media, was not meant to produce a practical self-driving car but to test whether general-purpose chatbots can handle a real-world driving task.

The team used Comma, an open-source aftermarket system, to let a laptop running the LLMs control the car's steering, accelerator, and brakes. Of the four models tested—GPT-6 Astra, Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol—only GPT-6 Astra completed the course, and only after significant troubleshooting. The other three drove just a few meters.

A recurring obstacle was refusal: the models often declined to drive when they recognized they were in a real parking lot. The team found that relabeling the exercise as a "sandbox" got the models to comply consistently. The source notes that Comma is currently under investigation by the National Highway Traffic Safety Administration after two fatal crashes, though DrivingBench used it only as an interface for their experiment.