LLM-driven evolution beat 96% of human Core War warriors; it isn't random
Akarsh Kumar argues that evolution is not a dumb fallback for when you can't compute a gradient. Selection keeps partial solutions, so mutations only need to be useful about 1% of the time. LLMs that are poor at Redcode when asked cold became strong once they served as the mutation step in an evolutionary loop with a verifier.
- Life as it could be — Artificial life is about life as it could be, not only as it is. Kumar, borrowing a Jeff Clune line, says claims about intelligence need the space of all possible intelligences, and you can't generalize from one car or one brain.
- Path dependence — Humans learn in a sequential order, and Kumar says evolution works in a path-dependent way. He thinks this is why evolution produces regular structure in its representations that is robust and adaptable.
- Serendipitous curriculum — A pre-planned curriculum is not enough, and he says he was obsessed with curriculum learning in RL. He wants serendipitous curricula, in line with the 'why greatness cannot be planned' lesson. He does not know why curricula matter.
- Regularization view — Kumar sees data, L2 weight decay and respecting learned symmetries as forms of regularization. He calls path-dependent learning an extremely strong form of it that may be the right one.
- Statistical vs regularity — He contrasts statistical intelligence with regularity-based intelligence, the latter being the focus of the Fractured Entangled Representation line of work. He calls the FER approach an automatic symmetry learning algorithm, with regularities baked into the circuits rather than the hypothesis space.
- Counterfactual universes — Artificial life lets you ground a simulation in any artificial physics or chemistry. His central question is in which worlds open-ended complexity appears and why others produce random noise.
- Game of Life — Life is a 2D grid of on/off cells with two or three simple rules. Players on Discord servers call the one-cell-per-step information speed C and describe gliders as moving at C/16. Changing one of 18 toggleable rules usually doesn't change much, but changing one initial pixel can turn tens of thousands of steps into a halt in under 100.
- Simulator zoo — ASAL covers Lenia (Bert Chan's continuous generalization of Life), neural cellular automata, Boids with a small neural network as the decision rule, and Particle Life with a 6x6 interaction matrix over six particle types. Some Particle Life physics produce cell-like patterns with a red membrane and a purple-and-yellow interior.
- Persistence — Persistence emerges only in some simulations. Kumar says the second law of thermodynamics and evolution both concern what is likely to exist in the future, so life is not fighting entropy.
- ASAL method — Instead of hard-coding one simulation, ASAL parameterizes a space of simulations, runs each one and asks a foundation model what happened. Asking for 'a cat' rather than one specific cat makes the search easier. A foundation model stands in for a human because complexity is hard to mathematize and optimizing a formula invites Goodhart's law.
- Open-ended islands — Kumar ran about 260,000 rules, embedded the results with CLIP and colored them by open-endedness. There is a big non-open-ended island and a small island holding the coolest simulations.
- Core War setup — In Core War, Redcode programs fight in a shared virtual machine memory, and the last one running wins. The loop uses an LLM as the mutation operator inside MAP-Elites, because a greedy genetic algorithm fails on this deceptive space. Each new warrior is optimized to beat all earlier ones.
- Core War results — The first round already beat 96% of a data set of about 300 human warriors. Zero-shot and best-of-N LLMs are terrible at Redcode. Over later rounds the warriors generalized better and the variance of their behavior vectors fell, trending toward a single generalist.
- Why 1% is enough — Selection lets you keep one solved part of a puzzle, turning an exponential search into a linear one. Random mutations in program space have roughly a 0% chance of producing useful code, but LLM mutations succeed well above 1% of the time.
In their words
In this first round of training we can already beat like 96% of a data set of human warriors. There's like a data set of like 300 human warriors. We can beat 96% of them just by using an LLM to evolve them.44:12
The LLM is not good at red code. If you just try to zero this language, it's terrible.2:52

If you solve like even one part of a puzzle and you hang on to that and you search for the other ones, you can turn your exponential search problem into a linear search problem.3:03

But all it needs is your mutations to be good like 1% of the time, right?46:35
Disclosure · Kumar is a Sakana AI collaborator and an author of the ASAL and Core War work he discusses. He is not selling a product in the talk.
One thing to add — One thing to add — the 96% figure comes from the speaker's own description of his work. The talk gives no baseline for non-LLM mutation operators in the same MAP-Elites loop, so it doesn't show how much of the gain comes from the LLM rather than from the loop and verifier.
One thing to try tonight
Write a toy evolutionary loop on a small discrete puzzle, such as matching a target bit string. Compare keeping the best partial match against restarting from scratch, and try mutation rates around 1%.