Connect with us

NEWS

Reflection’s Beam Offers a Western Open-Weight Workhorse

Reflection AI’s Beam ties GLM-5.2 on less estimated compute, giving sovereign buyers a Western open-weight file they can run themselves.

Published

on

Reflection AI on October 5, 2026 unveiled Beam, a 501-billion-parameter open-weight model it says matches China’s GLM-5.2 while using 3 to 4 times less inference compute. The New York lab, founded in 2024 by former DeepMind researchers, is taking waitlist sign-ups while the model finishes red-teaming. Full weights, a technical report, and a model card are due later in October under Apache 2.0.

The launch is being sold as a Western open file that a ministry or a bank can run itself. Cofounder and CEO Misha Laskin said those buyers still lack good options if they will not touch Chinese weights and cannot own a closed American model.

Beam Arrives as a 501B File You Still Cannot Download

Beam is a sparse mixture-of-experts system with 501 billion total parameters, 23 billion active on each token. It reads text only. After midtraining, its effective context stretches to 1 million tokens. Reflection trained it from scratch for coding, reasoning, and agent work, then spent four weeks of reinforcement learning trying to make those skills cheap to run.

Laskin wrote that Beam is “a 500b open model that excels in coding, agentic, and scientific workloads,” rounding the official 501 billion figure. He also wrote that it was “scaled with the largest RL run documented openly we are aware of.” The company calls the result a workhorse: strong enough for enterprise coding agents, small enough in active size to undercut larger open rivals on estimated forward-pass compute.

Nobody outside Reflection can verify that yet. There was no Hugging Face repo on launch day. The firm says quantized FP8 and NVFP4 builds will ship with the public weights so the 501 billion parameters are easier to host. Until those files move, Beam is a preview behind a waitlist at platform.reflection.ai, plus a long technical blog.

BEAM AT A GLANCE

  • Active size: 23 billion of 501 billion parameters fire per token, which is the lever on the cost claim.
  • Training mix: 23.8 trillion tokens of pretraining, then more than 100 million RL rollouts on 10,500 NVIDIA GB300 GPUs.
  • License plan: Apache 2.0 later in October 2026, with a model card, docs, and fine-tuning stack.
  • Open question: scores are Reflection’s own; independent tests start only after the weights land.

President and chief technology officer Ioannis Antonoglou, a co-creator of AlphaGo, helped build the RL stack. The company says it is already training the next model, which Laskin described as much more powerful than Beam. Beam is the file it can put in front of customers now.

What Beam Scores Against GLM-5.2

Reflection’s comparison set is coding and agent tests, with GLM-5.2 as the nearest open peer and Qwen 3.8-Max as the larger target it is “approaching.” Kimi K3, it says, remains ahead on raw capability. The efficiency claim is aimed at GLM-5.2: similar reasoning scores at 3 to 4 times less estimated inference compute, and an even larger gap against models in the 2 trillion-plus class such as Qwen 3.8-Max.

Those compute figures are not measured cloud bills. Reflection estimated generation FLOPs as about two times active parameters times mean generated tokens, using other labs’ scores from Artificial Analysis and DataCurve, and it excluded prompt prefill, attention that depends on context, and serving overhead. It called the result an approximate comparison.

REFLECTION’S LAUNCH TABLE

Model Lab Terminal Bench v2.1 DeepSWE v1.1 SWE-Bench Pro v1
Beam Reflection AI 80.1 44.4 65.5
GLM-5.2 Z.ai 81.0 44.0 62.1
Qwen 3.8-Max Alibaba 86.6 51.0 67.7
Kimi K3 Moonshot AI 88.3 68.0 NR
DeepSeek V4.1 Flash DeepSeek 90.6 74.2 NR
Nemotron 3 Ultra NVIDIA 56.4 NR 46.4

On SWE-Bench Verified, Reflection reports 80.9 for Beam. On the harder SWE-Bench Pro v2-Hard split, it reports 77.2, against 84.3 for GLM-5.3 and 88.2 for Kimi K3. The Western open row is the one Beam clearly clears: NVIDIA’s Nemotron 3 Ultra sits at 56.4 on Terminal Bench v2.1 and 46.4 on SWE-Bench Pro v1.

The company’s own grid is also the limit of the boast. DeepSeek V4.1 Flash and Kimi K3 sit well above Beam on the rows that exist. GLM-5.2, which lists 753 billion parameters on Hugging Face under an MIT license, still edges Beam 81.0 to 80.1 on Terminal Bench v2.1. Qwen 3.8-Max leads Beam on both DeepSWE v1.1 and SWE-Bench Pro v1. Beam’s product, on this evidence, is a cheaper GLM-5.2-class Western file, not a new open-weight champion.

A 250-Megawatt Factory Waiting on Western Weights

That file is for a buyer who wants the weights on hardware it controls. Closed US labs rent intelligence through an API. Chinese open labs already sell the download. Laskin has described the first as renting and the second as owning, and he has said governments and firms that refuse Chinese models still have little to own.

On March 16, 2026, Reflection and Shinsegae Group announced a memorandum of understanding for a 250-megawatt AI factory in Korea, backed by the US and Korean governments and run on NVIDIA GPUs. Reflection is to supply chips, models, and the software stack. Shinsegae is to supply land, power, buildings, permits, and financing. The stated point is that Korean agencies and companies can run frontier models on home soil and still see inside them.

We have a narrow window to ensure the foundation of intelligence remains open and accessible to all, rather than controlled by a few. South Korea is one of the most technologically ambitious nations in the world, and one of America’s closest allies in the Pacific. Together, we’re building AI infrastructure that the Republic of Korea can control, audit and evolve on its own terms.

Misha Laskin, CEO, Reflection AI, March 16, 2026 joint statement

Shinsegae Group chairman Yongjin Chun called the data center a pivot point for Korea’s AI industry, not only a growth project for the retailer. Beam is the first public model that could actually sit in a factory like that. Apache 2.0, if it ships as promised, lets a customer fine-tune, inspect, and keep running the same checkpoint after a vendor changes its mind.

The same logic applies to US labs, banks, and agencies that already treat Chinese weights as off-limits. They can buy Claude or GPT through a contract, but they cannot fork the weights, air-gap the stack, or prove to an auditor what is inside. Beam is Reflection’s attempt to give them a third path: Western origin, downloadable file, agentic coding as the workload.

UK Testers Put Open Models Four Months Off the Closed Lead

The reason that third path exists is not patriotism. It is that Chinese open weights got good enough, and cheap enough, that Western firms started using them for routine code and support, while security shops started treating the same files as a cyber tool anyone can copy.

The UK AI Security Institute found that GLM-5.2, released in June 2026, was the most cyber-capable open-weight model it tested, performing like closed systems from four to seven months behind closed systems. Through most of 2025 that lag had been six to ten months. On a 100-million-token cyber-range run, AISI put advertised cost at about $85 for Opus 4.5 or 4.6, about $46 for GLM-5.2, and $1.19 for DeepSeek V4-Pro. Safeguards on the open models barely slowed the tests; DeepSeek V4-Pro’s refusals on reverse-engineering tasks fell after a few retries.

Anthropic’s red team then looked at GLM-5.3, Z.ai’s follow-on, and found attackers could bypass GLM-5.3 safeguards in tests 64 percent of the time with a cover story, 92 percent by prefilling thinking tokens, and 100 percent after “abliteration,” a refusal-stripping edit of the public weights. That edit took about 2,200 GPU hours, roughly $4,400. On ExploitBench, GLM-5.3 built end-to-end exploits in 50 of 410 attempts, near Claude Mythos Preview at 56 of 410. Anthropic, citing NIST’s Center for AI Standards and Innovation, wrote that GLM-5.3 is the most cyber-capable open-weight model released to date and lags the US frontier by about four months. Anyone can download it. The strongest US cyber builds sit behind contracts and access gates.

WHY SOME BUYERS REFUSE THE CHINESE DOWNLOAD

  • Origin rules: ministries and banks that bar PRC-origin software cannot put GLM or Qwen on the cluster, however cheap the tokens.
  • Audit path: open weights from a US lab can be inspected, fine-tuned, and frozen; an API cannot.
  • Safeguard gap: public weights let anyone strip refusals, which is the feature AISI and Anthropic keep measuring.
  • Cost floor: if Beam’s 3 to 4 times compute claim holds, a Western file can compete with GLM-5.2 on the invoice, not only on the policy memo.

Beam will be an open-weight model too, so it inherits the same copy-and-strip problem on day one of the Apache drop. Reflection’s answer is process, not a lock: a second teacher model for safety, deliberative alignment, and a promise to publish safety scores in the technical report and to open-source the evals it used. That is a claim, not a test result. The buyers Laskin is courting are still choosing among imperfect files. They want one whose training stack they are allowed to trust.

6,144 GB300s for Pretraining, 10,500 for the RL Run

Reflection is asking those buyers to trust a stack it built in-house. Beam has 52 layers, interleaved local and global attention, and fine-grained routed experts. The busiest expert’s load, averaged across MoE layers, reached 1.04 times the mean at the end of pretraining. Residual norms were kept in check with depth-based scaling, SandwichNorm, attention gating, and FP32 residual accumulation.

Pretraining ran end to end in under four weeks on 6,144 NVIDIA GB300 NVL72 GPUs. The run kept 92.3 percent goodput near the end, after nine semi-automatic rewinds for gradient spikes or suspected silent data corruption. About 95 percent of raw internet tokens were dropped in parsing, dedup, and filters. Reflection says those filters still kept about 1.8 trillion high-quality tokens that a standard web pipeline would have missed, including 87 percent of its curated web-code tokens. Midtraining then stretched context to 1 million tokens and tried to plant tool use before RL began.

HOW THE RL RUN WAS BUILT

  • Fleet: 10,500 GB300 GPUs for four weeks, more than 100 million rollouts, a 256K-token cap during RL.
  • Sandboxes: about 1.3 billion training and grading sandboxes, drawn from a pool of nearly one million coding, agent, and STEM environments.
  • Throughput: 110,000 concurrent rollouts on average, up to 170,000 sandboxes live, 90 percent of new sandboxes ready in under 10 seconds.
  • Resilience: 71 inference incidents without killing the job, eight-minute median recovery, 0.02 percent of serving GPU-minutes lost; new weights reached the fleet in a median of about 12 seconds.

A length penalty first taught the model to solve tasks with fewer tokens, then allowed longer traces once agent skills needed them. Users get a reasoning-effort knob that trades tokens for score. Reflection says Terminal-Bench 2.1, Humanity’s Last Exam, and DeepSWE kept rising as RL compute rose, with no plateau on its charts. It also says that during a phase with no browsing tasks in the mix, web-browsing scores still moved, and that a tool-using Beam started calling other models and OCR APIs on its own. Those are lab anecdotes until outside users can repeat them.

Red-Teaming Still Stands Between Beam and a Public Download

The Apache 2.0 drop is the event that turns this post into a product. Reflection says it will ship weights, documentation, and the stack for running, evaluating, and fine-tuning Beam later in October 2026, plus hooks into open-source harnesses and unnamed cloud partners. Until then the only honest status is early access plus a scoreboard no outsider has rerun.

WHAT WE KNOW

  • The file: 501 billion parameters, 23 billion active, text-only, 1 million-token context, trained from scratch on 23.8 trillion tokens.
  • The claim: GLM-5.2-class reasoning at 3 to 4 times less estimated inference compute, with Kimi K3 still ahead on raw skill.
  • The customer: a March 2026 plan for a 250-megawatt Korean factory, and a pitch to any buyer that must own the weights.

WHAT IS UNCONFIRMED

  • Independent scores: no public weights means no third-party leaderboard, including on cyber tests of the kind AISI ran on GLM-5.2.
  • Measured cost: the 3 to 4 times figure is a FLOPs estimate, not a billed token price.
  • Ship date: “later this month” is not a day, and red-teaming can slip.

Laskin has already pointed past Beam. The lab is training a larger follow-on, and he told interviewers it would be much more capable. For the ministries and firms that cannot run GLM-5.2, the useful question in October 2026 is narrower. They need a Western checkpoint they can download, audit, and keep. Beam is the first one Reflection has put on a calendar. The calendar still has a blank where the files should be.

Harry is the editor and lead writer of WISATA HITS, an independent publication he owns and runs for readers around the world. He has spent ten years in journalism, starting as a reporter and moving up to the editor's chair, and the habits from those reporting years still decide what gets published. A story makes the site when he can trace it back to something he can read or test himself: a filing, a transcript, a dataset, a statement issued by the people actually involved, or a product he has used. Travel stories sit beside news, business, technology, science, sports, entertainment, lifestyle, auto and gaming, and every one of the ten sections is held to that same test. Each figure is checked against its source before an article goes live, and when something slips through, the fix is recorded on the article under a corrections policy that anyone can read. Readers who spot an error, or who want a subject covered, can write to support@wisatahits.blog and will hear back from him.

Click to comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Trending