Feature

Beam took 10,000 GB300s for RL, more compute than pre-training

Source · Beam: The Great American Open Model with ReflectionAI Co-Founder and CEO Misha Laskin
No Priors · Misha Laskin · Reflection AI · 2026-10-09 uploaded · 70min

한국어·English

Reflection AI CEO Misha Laskin says the 500B-parameter open model Beam spent more flops on reinforcement learning (10,000+ GB300s for four weeks) than on pre-training (6,000 GB300s). He also says open models now take 70% of tokens on gateways like OpenRouter and Vercel, the reverse of six months ago, and that enterprises must pass through a "rental" stage before they own their intelligence.

  • Team growth — Reflection went from about 30 people a year ago to around 300, and has now released Beam, its first open model.
  • Why build models — The company started 2.5 years ago as an RL bet on a ready-made open base model; a year in, the good open models were all from China, and Laskin found that RL at scale needs you to pre-train your own model.
  • Cost to the frontier — Catching the frontier cost hundreds of millions of dollars 18 months ago, single-digit billions about six months ago, and he expects order tens of billions next year, with roughly a 4x compute multiplier per model generation.
  • Efficiency gains — Human researchers alone gain about 7x compute efficiency per year; with models helping, he puts RL efficiency gains near 30x, at least four times faster than researchers alone.
  • Beam training run — Beam is 500B parameters total with 23B active; pre-training used 6,000 GB300s for a few weeks (about 12 days now), and RL used over 10,000 GB300s for four weeks.
  • RL never plateaus — Reflection's RL curves keep rising, so he frames further progress as an economics decision about how much to spend, not a question of whether training works.
  • Target domains — Code and agentic work is the foundation; the model is jagged across harnesses (e.g. Terminal Bench versions) but adapts quickly with new data, and the named enterprise verticals are finance KYC and compliance, cyber defense, and legal.
  • Reasoning efficiency — Beam is 3 to 4 times more efficient than models of the same capability class, and about 10x against some larger models, which he attributes to a strong pre-trained base plus the largest open-source RL run he has seen documented.
  • Rent vs own — Reflection monetizes by selling enterprises and sovereigns the cluster management, inference software, harness and services around an open model, so they own intelligence instead of renting tokens.
  • Token share — On gateways like OpenRouter and Vercel, token share flipped from 70/30 closed-to-open about six months ago to 70/30 open-to-closed, and he expects an operating-system-like split where Linux-style open software runs 95%+ of servers.
  • Enterprise path — He expects enterprises to customize systems (harnesses) before fine-tuning, and to move to open models only after ramping workloads on closed ones, typically at spend of $100 million or more a year.
  • Win formula — His formula for durable revenue is intelligence density times compute times customer trust, and he says AI-native companies already negotiate closed-model deals harder because open alternatives exist.
  • China's open models — He calls China's open models a gift to the West, says Chinese labs distill closed models at industrial scale, and argues they stay open because open models are geopolitical Trojan horses for infrastructure such as Huawei chips.
  • Safety argument — He argues openness is the default route to safety, citing 1990s encryption and Linus's law, and says removing cyber offense capabilities also removes defense, pointing to an OpenAI/Hugging Face incident where open models were used because guardrails blocked closed ones.
  • Team size — Frontier projects need around 100 researchers (AlphaGo had about 10), and he expects a steady state near that order, with growth in deployed engineers who build evals and harnesses for each enterprise capability.

In their words

to train beam which is a 500 billion parameter model total 23B active it was 6,000 uh GB300's we ran it I think for for a few weeks but now you know like with infrastructure efficiencies we can do it in about 12 days maybe less12:14
Misha Laskin slide · Beam: The Great American Open Model with ReflectionAI Co-Fou 12:14
Misha Laskin slide · 12:14 · No Priors
So beam tends to be three to four times more efficient than uh models of the same you know capability class and uh much more efficient when it comes to even you know there are models that are larger out there uh where you know where the efficiency gains are then end up being something like 10x.20:13
Misha Laskin slide · Beam: The Great American Open Model with ReflectionAI Co-Fou 20:13
Misha Laskin slide · 20:13 · No Priors
it was majority closed uh minority open when you go to any gateway like open router or versell and it's flipped almost exactly from 7030 closed open to 7030 open close.25:07
Misha Laskin slide · Beam: The Great American Open Model with ReflectionAI Co-Fou 25:07
Misha Laskin slide · 25:07 · No Priors
most the enterprise we talk to are considering now is uh by the time they've ramped up some significant workloads with closed models29:05
Misha Laskin slide · Beam: The Great American Open Model with ReflectionAI Co-Fou 29:05
Misha Laskin slide · 29:05 · No Priors
사전학습 GB300 6000개 강화학습 GB300(1 10000개
Beam 학습 GPU 규모 비교 — Misha Laskin이 Beam 학습에 쓴 GB300 수를 직접 언급(강화학습은 '10,000개 넘게', 4주).
18개월 전(억 달러대 1원문 단위 가늠 6개월 전(십억 달러대 10원문 단위 가늠 내년(백억 달러대) 100원문 단위 가늠
연도별 프런티어 추격 비용 — Laskin이 말한 자릿수(수억→수십억→수백억 달러)를 상대 규모로 환산한 것.

Disclosure · Misha Laskin is co-founder and CEO of Reflection AI, which makes Beam and sells deployment tools and services around open models.

One thing to add — One thing to add — The 70/30 token-share flip and the 30x RL efficiency figure are Laskin's own statements in the interview, and he gave no data or method for either. The Beam claims (3–4x efficiency, 10x versus larger models) also come from the company, and the tech report he mentions would be the place to check them.

One thing to try tonight
Pick one agent task you run on a closed model, log the number of steps and tokens it takes to finish, then run the same task on an open-weight model and compare the totals. Laskin says efficiency per task, not just capability, is what changes cost.