Reflection has come out of stealth with Beam, a text-only mixture-of-experts model totalling 501B parameters with 23B active. Trained from scratch, the model is aimed at coding, agentic and scientific work, and its full weights are due under Apache 2.0 this month. The company says pretraining used 23.8T tokens, partly from an OCR pipeline over hundreds of millions of PDFs, followed by a stable RL run on roughly 10,000 GB300s with over 100M rollouts across about 1M tasks.

Reflection's claimed results include 80.9 on SWE-bench Verified and 3–4x the inference efficiency of GLM 5.2. Independent observers, however, are more measured. One analysis estimates only ~12% BF16 MFU in pretraining, and another reads the architecture as an iso-FLOP replication of DeepSeek V3. Commentators place Beam around the GLM-5.2 level and below DeepSeek V4.1 Flash on some benchmarks.

The launch is notable less for topping leaderboards than for being a from-scratch US open-weight release. Nathan Lambert groups it with Nvidia and Thinking Machines as strong US releases that still trail Chinese counterparts. With full weights promised this month, Beam gives the US open-model ecosystem a new option, even if the frontier remains elsewhere.