OpenAI launched GPT-6.1 Sol at a significantly lower price point than its Astra model, claiming it beats GPT-6 Sol by 6.4 points on DeepSWE v1.1 and Opus 5.5 by 2.2 points on AutomationBench. Agent Arena data shows Sol Max entering at #5 with a median cost per task of $0.56, 39% cheaper than GPT-6 Sol while scoring 1.52 points higher.
Anthropic's Sonnet 5.5 debuted at #3 on Agent Arena and #1 in Chat, but its $2.74 cost per task is higher than Opus 5.5's $1.58, keeping it off the Pareto frontier. Anthropic models now hold the top three Agent Arena spots. Gemini 4 Argon took #1 in Text Arena.
The roundup also covers open models, decision models, and rumors. StepFun's Step 5 Preview ranks #7 among open-weight models, and llama.cpp added an endpoint for local decision-model inference. Perplexity claims its pplx-decider-v1-27b averages 85.7% across benchmarks, though skeptics call decision models a rebrand of zero-shot classifiers. Unconfirmed rumors mention a Claude Fable 5.5 and a delayed "Astra 6.1."