
What happened
Reflection AI started early access to Beam, its first open-weights model, with 501 billion total and 23 billion active parameters, pretrained on 23.8 trillion tokens.
Why it matters
Beam was built for coding, reasoning, and agent workloads, and Reflection AI says it matches or beats existing open base models of the same size, which could give developers a cheaper foundation model option.
What to watch
The efficiency claim hinges on external tests, since Reflection AI says Beam reaches scores comparable to larger open models at one-third to one-quarter the inference compute. Weights, a tech report, and model cards are due this month.
WHO IT HITSAI developers and engineering teams choosing among open-weights models for coding and agent workloads are the main audience, since Beam is positioned around cheaper inference rather than maximum scale. Enterprise teams evaluating self-hosted models may weigh it against existing open base models of the same size.
Summaries like this, in your inbox every morning.
Reflection AI is entering the open-weights market for the first time with Beam, a model it says was trained on 23.8 trillion tokens from web and proprietary licensed datasets. Rather than chasing the largest possible model, the company built Beam around coding, reasoning, and agent workloads, and says it matched or exceeded existing open base models of the same scale.
The training run itself was notable for its scale: Reflection AI says it generated more than 100 million rollouts across 10,500 NVIDIA GB300 GPUs over a four-week training period, and that it developed the algorithms, training environments, and infrastructure needed to sustain that high-compute reinforcement learning. That infrastructure claim is the foundation for the efficiency pitch that follows.
The efficiency pitch is where the competitive positioning lives. Reflection AI says Beam competes with large open models such as GLM-5.2 and approaches Qwen 3.8-Max on coding and agent tasks, while reaching comparable scores on advanced reasoning benchmarks at one-third to one-quarter the inference compute, with the gap widening against 2T+ parameter families. For developers weighing open models, the open question is whether independent testing confirms that efficiency advantage once the weights and technical report arrive this month, since the performance comparisons so far come from Reflection AI's own account.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Meta, Walmart, Stripe and Sierra Technologies are creating the Personal Agent Protocol, introduced by Sierra c…
Mistral AI opened a public preview of Mistral Large 4, its 1.05 trillion-parameter MoE model nicknamed "Le Cho…

Google began rolling out "Simple Guide" in Gemini Live on Android, letting users share camera or screen views…

Anthropic launched Claude for Google Workspace as a public beta for paid Claude plans, adding Claude to Google…

Testing a fake LLM call on CPython 3.14.8, the developer saw the first await succeed and the second raise Runt…

Anthropic released "Claude for Google Workspace" as a public beta, reachable from each tool's "Extensions(拡張機能…
