deniz.in

Markets

Weather

Loading weather

· via Hacker News – Front Page (hnrss.org)

Reflection launches Beam, a 501B-parameter open-weight model targeting frontier reasoning at lower cost

Reflection AI has unveiled Beam, a 501B-parameter open-weight MoE model that it claims matches leading Chinese open models on reasoning benchmarks while using 3-4x less inference compute. Weights are due this month.

Reflection launches Beam, a 501B-parameter open-weight model targeting frontier reasoning at lower cost

Reflection unveils its first open-weight model

Reflection AI, a two-year-old Brooklyn startup founded in 2024 by two former Google DeepMind researchers, has unveiled Beam, its first open-weight model. According to the company's announcement, Beam is a sparse mixture-of-experts model with 501 billion total parameters and 23 billion active per token, trained for coding, reasoning and agentic workloads. TechCrunch reports that the launch confirms earlier Axios reporting that the company was close to a release.

Beam is text-only. According to TechCrunch, it has a 1 million token context window and was pretrained on 23.8 trillion tokens from web and licensed proprietary datasets. For scale, TechCrunch compares it with Z.ai's GLM-5.2, at roughly 744 billion total parameters and 40 billion active — larger than Beam on both counts.

Efficiency is the headline claim

Reflection does not claim Beam is the strongest open model available. Its own write-up concedes that frontier open models such as Kimi K3 remain ahead on raw capability. The pitch is efficiency: Reflection says Beam scores on par with GLM-5.2 on advanced reasoning benchmarks while using 3-4x less inference compute, and approaches Qwen 3.8-Max on coding and agentic tasks even though models in that 2 trillion-plus parameter family need far more compute per token. The company describes Beam as a workhorse model for enterprise coding and agentic workloads.

TechCrunch adds that Reflection claims Beam outperforms today's leading Western open models, and that the startup's own benchmarks show it outscores Inkling — the open model from Mira Murati's Thinking Machines Lab released in July — on four coding tests where both report results, though Inkling is multimodal while Beam is text-only. TechCrunch notes the performance claims have not been independently verified.

Reflection also discloses a caveat: its FLOPs comparisons are estimated from active parameter counts and generated tokens, excluding prompt prefill, attention operations and serving overhead, so they approximate relative compute rather than measured inference cost.

A very large reinforcement learning run

Reflection credits two investments for Beam's capability: pretraining and high-compute reinforcement learning. The RL run lasted four weeks on 10,500 NVIDIA GB300 GPUs, generated more than 100 million rollouts at a maximum context length of 256,000 tokens, and used roughly 1.3 billion sandboxes for training and grading. To feed it, the company sourced nearly one million coding, agentic and STEM environments, primarily through synthetic data pipelines. Reflection believes this is one of the largest RL runs conducted by any open lab to date, and says benchmark scores kept improving as RL compute increased, with no sign of a plateau.

The announcement also describes engineering work on asynchronous policy gradients, where samples up to a full day old and 107 weight versions behind the current policy still produced stable learning. A controllable length penalty pushed the model to solve tasks with fewer tokens early in training, and a user-facing reasoning effort setting exposes the same tradeoff at inference time. Reflection reports transfer effects as well: during a training phase with no browsing tasks in the RL mix, browsing performance improved, and given web access the model learned to search for and query other language models and to call OCR APIs.

Funding and the sovereign AI pitch

According to TechCrunch, citing PitchBook, Reflection has raised about $4.7 billion from backers including Nvidia, Sequoia Capital and Lightspeed Venture Partners, with its last round valuing the company at $25 billion pre-money. It has also signed compute deals worth more than $7 billion with SpaceX and Nebius to secure Nvidia GB300 chips through 2029.

The company is aiming Beam at enterprises and sovereign customers, pitching AI factories that let institutions build customized local systems by training Reflection's models on their own data. TechCrunch reports that Reflection has begun testing a sovereign AI factory partnership with Shinsegae Group in South Korea, and that Axios found hedge funds and trading firms among the interested buyers.

Why it matters

Western open-weight models have trailed Chinese labs such as DeepSeek, Qwen, Z.ai and Moonshot for some time. Beam is a heavily funded attempt to close that gap, with inference efficiency rather than peak capability as the differentiator — a distinction that matters for cost-sensitive enterprise and agentic deployments where token spend dominates. For now the claims rest on Reflection's own benchmarks: weights, a technical report, a model card and developer artifacts are promised later this month, along with distribution through hyperscalers and neoclouds. Independent evaluation will determine whether the efficiency story holds.

  • #open-weight
  • #mixture-of-experts
  • #reinforcement-learning
  • #large-language-models
  • #reflection-ai

Related posts