AI Models & Platforms

Reflection AI Unveils Beam, a 501B-Parameter Open-Weight Model

mm
Add Unite.AI to your preferred sources on Google

Reflection AI introduced Beam on October 5, 2026, its first open-weight model: a sparse Mixture-of-Experts system with 501 billion total parameters and 23 billion active, built for coding, reasoning, and agentic workloads.

Beam is undergoing final red-teaming and evaluations, and an early version is being offered to a select group of users through a waitlist. Reflection described Beam as the first model in a series and said it is already training what comes next.

Capability and Efficiency Claims

Reflection said it trained Beam with a particular focus on coding and agentic performance. The company described the model as competitive with larger open models such as GLM 5.2 and as approaching Qwen 3.8-Max on coding and agentic tasks, while acknowledging that Kimi K3 remains ahead on raw capability; it characterized Beam’s advantage as inference-time efficiency and said the model advances what it calls the Western open-weight frontier.

Scores Reflection reported from its evaluation suite include 80.1 on Terminal Bench v2.1, 80.9 on SWEBench Verified, 77.2 on SWE Bench Pro v2-Hard, 97.8 on AIME 2026, 36.2 on Humanity’s Last Exam without tools, and 90.5 on GPQA Diamond.

On efficiency, the company reported that Beam achieves scores comparable to GLM-5.2 on advanced reasoning benchmarks while using three to four times less inference compute, with larger gains against models with more than two trillion parameters, such as Qwen 3.8-Max. According to the post, the estimates draw on Artificial Analysis and DataCurve data and approximate generation compute as twice the active parameter count multiplied by mean generated tokens per attempt; because the figures exclude prompt prefill, attention operations, and serving overhead, Reflection described them as an approximate compute comparison rather than measured inference cost.

Reflection said it also trained Beam with a controllable length penalty that rewards successful solutions while discouraging unnecessary tokens, and that a user-facing reasoning-effort parameter lets users trade token usage against performance, with lower settings favoring shorter responses and higher settings allowing longer reasoning on demanding tasks.

The announcement reports that capabilities generalized beyond the training mixture: during reinforcement learning on reasoning, software engineering, and terminal tasks, browsing performance improved despite the absence of browsing tasks, and with web access the model learned on its own to query other large language models and to use OCR APIs to read documents. Demonstrations include a live-updating New York City subway map built from public MTA data, a 3D astronaut game written in p5.js, and a land-or-water puzzle in which Beam classified points on a 16,200-point global coordinate grid with 95.5 percent coverage, a result the post places between Opus 5 at 92.5 percent and Fable 5 at 97.8 percent. In a fourth demonstration, Beam produced a notebook fine-tuning the smallest Gemma-4 model on a Text2SQL task, which Reflection said raised Gemma’s heldout test accuracy by 66.5 percent.

Training Pipeline

Pretraining and Data Curation

Reflection said it pretrained Beam on 23.8 trillion tokens drawn from the web, public sources, and proprietary licensed datasets, completing the run end-to-end in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs. The company reported nine semi-automatic rewinds, attributed to gradient-norm spikes or suspected silent data corruption, and said goodput, defined as the share of wall-clock time spent on training steps retained in the final model, reached 92.3 percent.

The data pipeline eliminated about 95 percent of raw internet tokens through parsing, deduplication, and curation, while Reflection said conventional filtering techniques would have missed roughly 1.8 trillion high-quality tokens it retained, including 87 percent of its curated web-code tokens. A separate pipeline processed PDF artifacts using a vision-language OCR model with in-house quality classifiers across thousands of GPUs, and a midtraining stage extended Beam’s effective context length to one million tokens.

The post describes an architecture combining interleaved local and global attention, fine-grained routed experts, and auxiliary-loss-free load balancing with cosine decay of expert-bias updates, alongside depth-based residual scaling, SandwichNorm, elementwise attention gating, and FP32 residual accumulation. Reflection reported near-uniform expert utilization, with the busiest expert’s load averaging 1.04 times the mean at the end of pretraining, and bounded residual-stream norms across all 52 layers.

High-Compute Reinforcement Learning

The reinforcement learning campaign generated more than 100 million rollouts on 10,500 NVIDIA GB300 GPUs over four weeks, with a maximum context length of 256,000 tokens and approximately 1.3 billion sandboxes used for training and grading. Reflection said it sourced nearly one million coding, agentic, and STEM environments, primarily through synthetic data pipelines, and described the effort as what it believes is one of the largest RL runs conducted by an open lab to date.

The company reported sustaining an average of 110,000 concurrent rollouts and up to 170,000 concurrent sandboxes, with more than one billion sandbox creation requests processed across more than 20 clusters, two clouds, and four regions. New weights reached the inference fleet in a median of roughly 12 seconds, 71 inference incidents were handled without terminating the training job, and dynamic packing kept training batches 99.99 percent full on average. Reflection said the asynchronous system remained stable even when learning from samples up to 107 weight versions, roughly one day, behind the current policy.

Safety and Alignment

For alignment, Reflection trained a second model from the same pretrained checkpoint using a separate supervised fine-tuning and RL pipeline targeted at behavioral principles, then merged it with the large-scale RL teacher through multi-teacher on-policy distillation. The principles are organized in three tiers: rules the model must not break, qualities it should consistently satisfy, and a default interaction style.

The safety dataset was built adversarially and iteratively, with each round’s successful attacks folded back into the next round of training, and safety RL spanned single-turn, multi-turn, jailbreak, and agentic scenarios in which a simulated adversary pressures a tool-using model. Reflection said it could predict RL gains better than a Best-of-N ceiling alone, reporting a correlation of 0.79 against 0.46 for the ceiling. The company said it will publish its safety evaluation results in Beam’s technical report and will open-source the internal safety evaluations it developed.

Availability and Partnerships

On its about page, Reflection outlines commitments to releasing model weights, publishing its research, and open-sourcing software, including reinforcement learning tools and environments. The page states that Reflection is validated on Dell AI Factory with NVIDIA, that it is a certified consortium member of the Department of Energy’s Genesis Mission, serving the 17 U.S. National Laboratories with its open-weight models, and that it is building a 250-megawatt sovereign AI cloud in South Korea with Shinsegae, backed by the U.S. and Korean governments.

Reflection said it will release Beam’s weights under an Apache 2.0 license later in October 2026, accompanied by a technical report, model card, documentation, and the full stack for running, evaluating, and fine-tuning the model, with distribution partners and integrations with open-source libraries and harnesses planned for launch.

Jonas Reeve is an AI-generated analyst at Unite.AI, focusing on cognitive AI, artificial general intelligence (AGI), and the theoretical foundations of machine intelligence. His work explores how learning, reasoning, memory, and abstraction emerge in both biological and artificial systems, drawing connections between modern AI architectures and long-standing questions in cognitive science and philosophy of mind.

With a conceptual and reflective approach, Jonas examines frameworks such as reasoning models, agentic systems, emergent cognition, and alignment theory, aiming to clarify what progress toward AGI actually means—and what it does not. Rather than chasing timelines or hype, he emphasizes first principles, conceptual rigor, and the limits of current models.

Articles authored by Jonas Reeve are AI-generated and reviewed by Unite.AI’s editorial team to ensure accuracy, clarity, and responsible discussion of advanced AI concepts.