Beam took 10,000 GB300s for RL, more compute than pre-training
Reflection AI CEO Misha Laskin says the 500B-parameter open model Beam spent more flops on reinforcement learning (10,000+ GB300s for four weeks) than on pre-training (6,000 GB300s). He also says open models now take 70% of tokens on gateways like OpenRouter and Vercel, the reverse of six months ago, and that enterprises must pass through a "rental" stage before they own their intelligence.
- Team growth — Reflection went from about 30 people a year ago to around 300, and has now released Beam, its first open model.
- Why build models — The company started 2.5 years ago as an RL bet on a ready-made open base model; a year in, the good open models were all from China, and Laskin found that RL at scale needs you to pre-train your own model.
- Cost to the frontier — Catching the frontier cost hundreds of millions of dollars 18 months ago, single-digit billions about six months ago, and he expects order tens of billions next year, with roughly a 4x compute multiplier per model generation.
- Efficiency gains — Human researchers alone gain about 7x compute efficiency per year; with models helping, he puts RL efficiency gains near 30x, at least four times faster than researchers alone.
- Beam training run — Beam is 500B parameters total with 23B active; pre-training used 6,000 GB300s for a few weeks (about 12 days now), and RL used over 10,000 GB300s for four weeks.
- RL never plateaus — Reflection's RL curves keep rising, so he frames further progress as an economics decision about how much to spend, not a question of whether training works.
- Target domains — Code and agentic work is the foundation; the model is jagged across harnesses (e.g. Terminal Bench versions) but adapts quickly with new data, and the named enterprise verticals are finance KYC and compliance, cyber defense, and legal.
- Reasoning efficiency — Beam is 3 to 4 times more efficient than models of the same capability class, and about 10x against some larger models, which he attributes to a strong pre-trained base plus the largest open-source RL run he has seen documented.
- Rent vs own — Reflection monetizes by selling enterprises and sovereigns the cluster management, inference software, harness and services around an open model, so they own intelligence instead of renting tokens.
- Token share — On gateways like OpenRouter and Vercel, token share flipped from 70/30 closed-to-open about six months ago to 70/30 open-to-closed, and he expects an operating-system-like split where Linux-style open software runs 95%+ of servers.
- Enterprise path — He expects enterprises to customize systems (harnesses) before fine-tuning, and to move to open models only after ramping workloads on closed ones, typically at spend of $100 million or more a year.
- Win formula — His formula for durable revenue is intelligence density times compute times customer trust, and he says AI-native companies already negotiate closed-model deals harder because open alternatives exist.
- China's open models — He calls China's open models a gift to the West, says Chinese labs distill closed models at industrial scale, and argues they stay open because open models are geopolitical Trojan horses for infrastructure such as Huawei chips.
- Safety argument — He argues openness is the default route to safety, citing 1990s encryption and Linus's law, and says removing cyber offense capabilities also removes defense, pointing to an OpenAI/Hugging Face incident where open models were used because guardrails blocked closed ones.
- Team size — Frontier projects need around 100 researchers (AlphaGo had about 10), and he expects a steady state near that order, with growth in deployed engineers who build evals and harnesses for each enterprise capability.
In their words
to train beam which is a 500 billion parameter model total 23B active it was 6,000 uh GB300's we ran it I think for for a few weeks but now you know like with infrastructure efficiencies we can do it in about 12 days maybe less12:14

So beam tends to be three to four times more efficient than uh models of the same you know capability class and uh much more efficient when it comes to even you know there are models that are larger out there uh where you know where the efficiency gains are then end up being something like 10x.20:13

it was majority closed uh minority open when you go to any gateway like open router or versell and it's flipped almost exactly from 7030 closed open to 7030 open close.25:07

most the enterprise we talk to are considering now is uh by the time they've ramped up some significant workloads with closed models29:05

Disclosure · Misha Laskin is co-founder and CEO of Reflection AI, which makes Beam and sells deployment tools and services around open models.
One thing to add — One thing to add — The 70/30 token-share flip and the 30x RL efficiency figure are Laskin's own statements in the interview, and he gave no data or method for either. The Beam claims (3–4x efficiency, 10x versus larger models) also come from the company, and the tech report he mentions would be the place to check them.
One thing to try tonight
Pick one agent task you run on a closed model, log the number of steps and tokens it takes to finish, then run the same task on an open-weight model and compare the totals. Laskin says efficiency per task, not just capability, is what changes cost.