Founders say LLM latency is halving monthly, enabling robot control
Two startup founders argue the path to general-purpose robots runs through general-purpose language models, not robot-specific training. They say frontier labs and robotics companies now expect human-like general-purpose robots within two years, a timeline they argue society isn't prepared for.
- RT2 origins — Jay traces current robot-use agents back to Google's RT2 paper, which fine-tuned a web-pretrained language model to output end-effector poses instead of English text.
- Bitter lesson framing — Jaime and Hamming argue the lesson isn't architecture but data: pouring coding, computer-use and egocentric data into models produces general-purpose agents that outperform robotics-specific models.
- Code as policies — The founders cite Google DeepMind's 'code as policies' work and the Voyager Minecraft agent as early proof that coding agents could write robot-control code one-shot without robot-specific data.
- In-context limits — Francois describes an unpublished experiment showing in-context learning on held-out tasks like GSM8K improves non-monotonically and saturates after roughly 20-40 examples, capping near half the model's trained context length.
- Harness design — Waddle Labs packages in-context learned skills into reusable code or tool calls so a smaller, faster model can run repetitive tasks instead of keeping a large model like Astra in the loop at every step.
- Live demo — In a demo, the model Astra controls robot arms via camera feeds to pick a block and place it in a bowl, issuing tool calls turn-by-turn rather than full code policies.
- Latency trend — Jaime says stable-class LLM latency is improving roughly 2x per month, and if the trend continues real-time robot control could arrive by the end of the year.
- Platonic hypothesis — Citing Philip Isola's Platonic Representation Hypothesis, Jay argues a sufficiently strong language model and a strong robotics model should converge to similar world representations, meaning one strong general model could beat specialized ones.
- Why Astra improved — The founders attribute Astra's spatial-intelligence gains to heavy pretraining on computer-use data, such as dragging a cursor to orbit CAD objects in Blender, which teaches spatial concepts like top-down and left-right.
- 2-year timeline — Jay says there is consensus among frontier labs and robotics companies that general-purpose robots capable of a 'competent teenager's' tasks from natural language instructions could arrive within two years or sooner.
- Sleep analogy — Hamming compares future training loops to DAgger-style data aggregation or a 'sleep phase,' where collected in-context experience is periodically distilled back into updated model weights, similar to the DreamCoder system's skill-library refactoring.
In their words
I feel like at the end of the day, it's a lot about the bit of less, right? If you give the agent or if you give the AI model more autonomy and you if you unshackle it a bit more and give it more resources, it can actually do a lot of things that we fine-tuned it to do.3:55

one one kind of data is now these models can uh well, I mean they they write code much better. So, now they can uh write complex policies as code.6:03

There's a few things that models have got much better at since RT2.5:52

So, one thing that we saw is that for stable class LLMs, their latency is improving by around 2x per month.17:24

Disclosure · The speakers are founders of Waddle Labs and RoboCurve, the two startups whose robot-control demos and eval work are the subject of the interview, and the episode is produced by Y Combinator, which funds startups including possibly these companies.
One thing to add — One thing to add — the 2x-per-month latency claim and the two-year general-purpose robot timeline are founder predictions stated as near-consensus, not independently verified benchmarks, so they should be read as directional rather than measured. The in-context learning saturation experiment Francois describes is also explicitly unpublished.
One thing to try tonight
Try prompting a coding-capable LLM (like the ones referenced, e.g. GPT-class models) to write "code as policies" for a toy robotic/simulated-arm task using a small set of Python functions (pick_up, move_to, lift) and see how far one-shot in-context reasoning gets without any task-specific fine-tuning.