Reflection AI’s Beam trails GLM-5.3 and Kimi K3 on its own benchmarks
Reflection AI launched Beam on 5 October, a 501 billion parameter open-weight model, and the company’s own benchmark table puts it level with Z.ai’s GLM-5.2 while trailing GLM-5.3, Kimi K3 and DeepSeek V4.1 Flash. Neither of the two trade write-ups of the launch reproduced that table, which matters because the blog post says the gap out loud.
Where frontier open models like Kimi K3 remain ahead on raw capability, Beam’s advantage is efficiency at inference time.
Reflection AI, Introducing Beam
What Reflection’s own table shows
Four rows from the figure in the announcement carry the point, so they are worth reading together. Beam sits within a point of GLM-5.2 on each of them, and below every newer Chinese open model that reported a score.
| Benchmark | Beam | GLM-5.2 | GLM-5.3 | Kimi K3 | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|
| DeepSWE v1.1 | 44.4 | 44.0 | 61.0 | 68.0 | 74.2 |
| Terminal Bench v2.1 | 80.1 | 81.0 | 88.2 | 88.3 | 90.6 |
| HLE, no tools | 36.2 | 40.5 | 42.3 | 46.9 | 39.1 |
| GPQA Diamond | 90.5 | 91.2 | 91.7 | 93.5 | 90.9 |
That 44.4 on DeepSWE is worth sitting with, because this site has reported the same benchmark before. Together AI measured GLM-5.3 at 69.0% pass@1 across 113 DeepSWE tasks, and noted the public leaderboard read 44% for the previous GLM generation. Reflection’s harness puts GLM-5.3 at 61.0 instead, so the two runs disagree, but both place Beam where GLM-5.2 was.
Beam still wins things. It beats GLM-5.2 on AutomationBench, 37.0 to 26.2, on MCP Atlas, 78.7 to 77.8, and on SWE Bench Pro v1, 65.5 to 62.1. Against Thinking Machines Lab’s Inkling it leads every coding row that both report, though Inkling edges it on the AA Omniscience public split, 14.2 to 13.0.
The size comparison is what makes the efficiency pitch land, though, because Beam is the smallest model in the group. Inkling, which we covered when Thinking Machines released it as open weights, carries 975 billion total parameters and 41 billion active, and handles images and audio. Kimi K3 is bigger still: Moonshot calls it the first open model to reach 2.8 trillion parameters. Beam is 501 billion total, 23 billion active, and text only.
The 3-4x efficiency claim is an estimate
Reflection says Beam scores comparably to GLM-5.2 while using “3-4x less inference compute”, a figure TechCrunch reported alongside the note that the performance claims haven’t been independently verified. The footnote under the efficiency chart goes further. Reflection estimated forward-pass compute as twice the active parameter count times mean generated tokens, excluding prompt prefill, attention operations and serving overhead.
In Reflection’s own words those numbers “represent an approximate compute comparison rather than measured inference cost”. That caveat didn’t make either write-up, and it’s the load-bearing one, since efficiency is the whole case for a smaller model.
The training run itself is less contested, and that is where the detail is firmest. Beam was pretrained end to end in under four weeks on a cluster of 6,144 Nvidia GB300 NVL72 GPUs, on 23.8 trillion tokens. The reinforcement learning phase then ran 10.5K GB300s for four weeks, generating more than 100 million rollouts and using roughly 1.3 billion sandboxes. SiliconANGLE rounded that fleet to 10,000 cards and called the model open-source.
Reflection calls it open-weight instead, and the weights aren’t out yet. The model is in final red-teaming behind an early access waitlist. Reflection says it’ll publish the weights under an Apache 2.0 license this month, with documentation, a model card and a technical report, a licence neither write-up named.
The compute deals behind it are reported two ways. SiliconANGLE describes a $6.3 billion arrangement with SpaceX for the GB300 NVL72 appliances used to train Beam, while TechCrunch counts deals worth more than $7 billion with SpaceX and Nebius running through 2029. Reflection’s post names neither partner, and the company didn’t respond to TechCrunch in time for publication.
But what’s checkable arrives this month. Once the Apache 2.0 weights and the technical report land, outside labs can run Beam on their own harnesses and measure inference cost rather than estimating it. Until then the strongest claim the table supports is parity with last generation at a fraction of the size.
Get the daily rundown
One email each weekday with the AI news that matters, every claim linked to its primary source.
Free, one email each weekday, unsubscribe in one click. We never sell or share your address.
