TL;DR
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Reflection has introduced Beam, its first open-weight model, with 501 billion total parameters and 23 billion active per token. The company reports strong coding and agentic benchmark results and lower inference-compute needs on some comparisons, but Beam is still undergoing red-teaming and evaluation; the weights and technical documentation have not yet been released.
Reflection has introduced Beam, its first open-weight model, a sparse mixture-of-experts system with 501 billion total parameters and 23 billion active parameters per token. The company says the model is aimed at coding, reasoning and agentic workloads, but its weights are not yet available: Beam is undergoing final red-teaming and evaluation, with release materials planned for later this month.
Reflection says Beam was pretrained on 23.8 trillion tokens from curated web material and proprietary licensed datasets. The company also reports a large reinforcement-learning campaign: more than 100 million rollouts generated over four weeks using 10,500 NVIDIA GB300 GPUs. It says training and grading used about 1.3 billion sandboxes and drew on one million coding, agentic and STEM environments.
In benchmark results published with the announcement, Reflection reports Beam scores of 80.1 on Terminal-Bench 2.1, 80.9 on SWE-bench Verified and 77.2 on SWE-bench Pro v2-Hard. These are company-reported results; the supplied announcement does not provide an independent evaluation of the figures. Reflection describes Beam as competitive with larger open models on coding and agentic tasks, while acknowledging that models including Kimi K3 remain ahead on raw capability.
Reflection also claims Beam can match GLM-5.2 on advanced reasoning benchmarks with three to four times less inference compute. Its comparisons estimate generation compute using active parameters and generated tokens, and exclude prompt prefill, some attention costs and serving overhead. They are estimates, not direct measurements of total production cost. The company says this efficiency could make Beam useful for enterprise coding and agentic workloads.
Lower Compute for Coding Tasks
For organizations running coding assistants or agents at scale, inference requirements can affect both operating cost and how many tasks a system can handle. If Beam’s reported efficiency holds in independent testing and real deployments, its lower estimated compute per response could make a capable model more practical to serve, even though its total parameter count is large.
The announcement also highlights the resources behind open-weight development: a four-week run on 10,500 GB300 GPUs, plus extensive environments for reinforcement learning. That scale may make replicating the training effort difficult for smaller teams. The practical significance for developers will depend on the eventual license, hardware needs, serving performance and whether the benchmark advantages persist outside the company’s own tests.
As an affiliate, we earn on qualifying purchases.
How Reflection Trained Beam
Beam is a sparse mixture-of-experts model: its stated total parameter count is 501 billion, while 23 billion are active per token. Total size and active size describe different aspects of the architecture; neither figure alone establishes the hardware required for a particular deployment. Reflection says it focused training on coding and agentic performance, combining large-scale pretraining with reinforcement learning.
The company says its RL approach used asynchronous policy gradients, where training can learn from rollouts produced by earlier model versions. Reflection reports developing methods to manage policy staleness and differences between training and inference systems. It says the run remained stable even with samples more than a day old, but detailed methods and supporting technical material have not yet been published.
Reflection frames Beam as a step forward for Western open-weight models, while its own comparisons place some competitors ahead on capability. Those assessments are the company’s characterization, not an independent ranking. The full technical report and model card are expected to provide more detail on evaluation methods, training and intended use.
“Beam is our first open-weight model.”
— Reflection
As an affiliate, we earn on qualifying purchases.
Testing and Release Details Pending
Beam has not yet been released, and Reflection says final red-teaming and evaluations are still underway. The announcement does not specify a release date, the terms of the model’s license, access requirements or the hardware needed to run it. Those details matter to developers deciding whether the model can be used, modified and deployed for their purposes.
The published benchmark results and compute comparisons are presented by Reflection, with some comparisons drawing on Artificial Analysis and DataCurve. Independent replication, full evaluation protocols and results across a broader range of tasks are not included in the supplied material. The compute estimates also exclude prompt prefill, context-dependent attention operations and serving overhead, so they should not be treated as complete measures of deployment cost.
Reflection says its RL campaign showed no sign of a performance plateau as compute increased, but the announcement does not establish how that trend will generalize to other training runs or tasks. The company also describes the run as one of the largest conducted by an open lab, qualifying that characterization as its belief rather than presenting it as independently verified.
GPU for machine learning inference
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Weights and Technical Report Due
Reflection says it plans to publish Beam’s weights, technical report, model card and developer artifacts later this month, after its current red-teaming and evaluation work. The company is also accepting sign-ups for early access. No exact publication date is specified in the announcement.
Those materials should clarify the model’s license, evaluation setup, safety testing, deployment requirements and the details behind its training and compute claims. Independent tests and developer experience after release will help establish whether Beam’s reported coding performance and inference efficiency translate into practical results.
large language model deployment hardware
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What is Beam?
Beam is Reflection’s first open-weight model, a sparse mixture-of-experts system designed for coding, reasoning and agentic workloads. Reflection lists 501 billion total parameters and 23 billion active parameters per token.
Can developers download Beam now?
No. Reflection says Beam is undergoing final red-teaming and evaluations. The company plans to release the weights and related materials later this month and offers an early-access sign-up.
What performance results has Reflection reported?
Reflection reports scores including 80.1 on Terminal-Bench 2.1, 80.9 on SWE-bench Verified and 77.2 on SWE-bench Pro v2-Hard. These are company-published results; independent confirmation is not included in the announcement.
What does Reflection mean by inference efficiency?
Reflection says Beam can achieve reasoning results comparable to GLM-5.2 using three to four times less inference compute. Its comparison is an estimate based on active parameters and generated tokens, and excludes several costs, including prompt prefill and serving overhead.
When will the model license and technical details be available?
Reflection says it plans to release a technical report, model card and developer artifacts with the weights later this month. The supplied announcement does not give a specific date or state the license terms.
Source: hn
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
