Beam Explained: 501 Billion Parameters, Only 23 Billion at Work

Reflection AI unveiled Beam on October 5, 2026, a 501-billion-parameter open-weight model built for coding, reasoning and agentic tasks. Open-weight means outside users can inspect and run the trained system. This is the company's first open model, according to Fortune, and the company detailed it on its research blog on the same date.
Beam uses a sparse Mixture-of-Experts design, which keeps many expert parts but calls on only a few per job. Total parameters stand at 501 billion. Active parameters per forward pass, or single calculation step, stand at 23 billion, as stated by Reflection AI. The design target is coding, reasoning and agentic workloads, which involve tool use, multi-step planning and code writing.
The company positioned Beam to rival Chinese open models at lower compute cost, according to TechCrunch. Reflection AI claims Beam is three to four times more efficient than rival open models from Western companies, as reported by Fortune. The company has not in the verified disclosures published the benchmark suite, token counts or hardware baseline behind that efficiency ratio in a form that permits independent replication from the facts provided.
Distribution follows a dual track. Reflection AI states it develops open-weight models that users can inspect, fine-tune and deploy. It states it serves open models through its API platform and supports deployment in private cloud, on-prem, air-gapped and edge environments. It also states it develops open-source software including deployable containers, customization recipes and agentic harnesses.
Reflection AI is backed by Nvidia. It raised $2 billion in a funding round that valued the company at $8 billion, as reported on October 9, 2025. The Wall Street Journal reported on March 26, 2026 that the company was in talks to raise $2.5 billion at a $25 billion valuation, according to Reuters. The later figure describes talks, not a closed round.
The broader context here is cost per useful token, not headline parameter count. For practitioners, a 501B/23B sparse configuration implies conditional compute. Memory footprint tracks total parameters. Inference FLOPs track active parameters plus routing overhead. If routing holds under agentic rollouts with long contexts and repeated tool calls, the operator saves compute. If load balancing degrades or experts co-activate, the saving compresses.
Looking at what this means for infrastructure buyers and model operators, deployment options carry as much weight as the weights themselves. API access shifts capacity planning to the vendor. Self-hosting in private cloud or air-gapped settings keeps data and customization in-house but leaves the operator with sharding, KV-cache management and expert parallelism. Containers and fine-tuning recipes lower integration cost, but they do not remove evaluation cost. Any claim of three to four times efficiency needs to be tested against specific concurrency, sequence length, quantization and quality thresholds before it enters a budget model. Until that testing is public, treat the ratio as a vendor claim to be verified, not a planning input.


