TypeSafe AI's Jev, a decision model that picks from a list of options rather than generating text or images, beat Pokémon Red in under a week, according to developer Andrew Boyd. The victory came on September 23, 2026, when Jev defeated the Elite Four and Champion to enter the Hall of Fame. This contrasts with traditional LLM-based chatbots that have taken months to complete similar Pokémon games, with Anthropic's Claude Plays Pokémon stream still unfinished as of January.

However, Jev didn't do it alone. Anthropic's Claude Opus 5 monitored the game log and acted as a coach, modifying the options and data available to Jev when it got stuck. The harness changelog had 474 entries, mostly failures paired with responses, such as walking into Lorelei's shut entrance 53 times or crossing a Rock Tunnel ladder 124 times in ten minutes. Human viewers also contributed by sending tips in chat, which were used to improve Jev's options.

A second Jev-based run by Christian Mathiesen at Frigade took a different approach, reading the game's memory to list legal options, but it never wrote to memory. Mathiesen noted his first version, where Jev could choose buttons directly, never left Pallet Town. A separate experiment by stmonty trained a small world model on an RTX 3080 Ti with 42,000 frames, but it only picked a starter in 52 of 100 tries, showing far less capability.

The success of Jev's run demonstrates that collaboration between specialized models can improve problem-solving, but it also underscores limitations. TypeSafe itself says open-ended tasks are better suited to an LLM, and the heavy reliance on Claude, the developer, and the audience suggests Jev's role was more about decision-making within a constrained framework than autonomous play. The sources agree on the outcome and the collaborative nature, though they differ on the specific approaches and costs, with Mathiesen estimating about $1–1.70 per 24 hours for his run. Overall, the achievement highlights the potential of hybrid AI systems, but also the need for human and LLM oversight in complex tasks.{