Technology

Reflection's Beam Challenges Chinese Open Models at Lower Compute Cost

Martin HollowayPublished 34m ago4 min readBased on 3 sources
Reading level
Reflection's Beam Challenges Chinese Open Models at Lower Compute Cost
source:reflection.ai

Reflection AI unveiled Beam, its first frontier open-weight AI model, on Oct. 5, 2026. TechCrunch

What Beam is

Beam is a text-only mixture-of-experts model, a design that holds many specialist subnetworks and switches on only part of them for each request. It was trained with high-compute reinforcement learning, learning from rewards for correct step-by-step behavior, for reasoning, coding and agentic tasks where software uses tools to complete work. Reflection

The text-only choice narrows the focus. Training concentrates on tokens, tool traces and code execution, rather than spreading capacity across vision or audio. Tokens are the short text chunks models read and write.

Beam has 501 billion total parameters and 23 billion active parameters, the weights used on any single pass. Pre-training covered 23.8 trillion tokens, and the model supports a 1 million token context window, the amount of text it can hold in mind at once. TechCrunch Only a small fraction of weights engages per forward pass, the computation that produces one output.

Reflection claims Beam scores on par with Z.ai's GLM-5.2 on advanced reasoning benchmarks while using 3-4x less inference compute, the computing needed to run the model for users. GLM-5.2 has roughly 744 billion total parameters and 40 billion active parameters. Reflection also claims Beam outperforms leading Western open models on advanced reasoning benchmarks. The claims center on inference efficiency, not just raw score. For deployment teams, that distinction matters more than leaderboard position.

For comparison, GLM-5.2 is available on Z.ai, and its model weights are publicly available on HuggingFace and ModelScope. Z.ai Both systems are accessible as open weights, which allows direct testing on internal reasoning and coding suites. That shared access is why the comparison matters to Western labs and enterprise builders.

How it was trained

Reflection built a pool of nearly one million environments for Beam's reinforcement learning training. Environments are practice tasks with checkable answers. They were created primarily through synthetic data pipelines, supplemented by proprietary vendor data and open-source sources. Reflection

Across the full campaign, Reflection generated over 100 million rollouts, trial runs where the model attempts a task and learns from the result. Eighty million of those rollouts trained Beam's reasoning expert alone. For comparison, Reflection states Inkling was trained on 30 million rollouts and MiMo on 753,000 rollouts. The jump points to environment diversity and rollout volume as the lever, rather than pre-training token count alone.

Reflection said it used Artificial Analysis and DataCurve as sources for other models' evaluations in its inference-efficiency comparisons. That detail is useful for practitioners who track third-party inference pricing and throughput. It does not resolve benchmark comparability on its own, but it makes the efficiency claim auditable against public evaluation infrastructure.

Funding and computing power

Reflection was founded in 2024 by two former Google DeepMind researchers. It has raised roughly $4.7 billion from backers including Nvidia, Sequoia Capital and Lightspeed Venture Partners, per PitchBook. That budget supports a 23.8 trillion token pre-train followed by a 100 million rollout reinforcement learning phase at a two-year-old lab.

Compute access was secured separately. Reflection signed deals collectively worth more than $7 billion with SpaceX and Nebius to secure access to Nvidia GB300 chips through 2029. That timeline extends well beyond Beam. It covers future training runs and inference capacity, and it ties accelerator supply to nontraditional cloud partners.

The broader context here is a shift in where open-model competition happens. Pre-training scale once dominated discussion. Active parameters and inference cost per reasoning task now shape adoption. A 23 billion active model that matches a 40 billion active model at lower compute would change serving economics for coding assistants, software agents and long-context workflows, where margin is set by tokens generated per dollar and latency under load.

In my view, the number worth watching is not 501 billion. It is nearly one million environments. Worth flagging, synthetic environment generation at that scale introduces its own failure modes around distribution skew and reward hacking, which only emerge under extended agentic rollouts. If Reflection has managed to keep environment quality high while pushing to 100 million rollouts, that pipeline is the durable asset. It would transfer to future models more directly than any single checkpoint, and it would give enterprise teams an open-weight option tuned for verifiable, multi-step work rather than single-turn response.