The Curious Mind: A Brief History Of Intelligence
What six hundred million years of brain evolution says about the machines we're building, and about us
"Nothing in biology makes sense except in the light of evolution."
- Dobzhansky
“What I cannot create, I do not understand.”
- Feynman
“We are not thinking machines that feel, we are feeling machines that think.”
- Damasio
"We tell ourselves stories in order to live."
- Joan Didion
“We don't see things as they are, we see them as we are.”
- Anais Nin
I majored in CompSci twenty seven years ago, but my first love was Child Psychology, I’ve always been fascinated by how the human brain learns. It is amazing how a baby is born, and it learns how to navigate the world in just a few months. I saw this with my own and its was beautiful.
Let me give you my biases, I don’t think our current course of AI progress will lead us to anything approximating human intelligence, because I don’t think we fully understand human intelligence. While we may get AGI - artificial general intelligence, it will be very artificial and hardly the general type we expect from humans.
Back to the subject of the day. I’m always on the look out for books that can help explain human intelligence, because understanding human’s is a better way to build smart machines because we are a working example of General Intelligence.
All of this drew me to Max Bennett’s A Brief History of Intelligence.
Max spent his career selling machine learning to companies. Recommendation engines, customer segmentation, the unglamorous middle of the industry. He kept hitting the same wall: the models were superb at things people find hard and useless at things a five-year-old finds obvious (Moravec's Paradox). So he started reading neuroscience in the evenings, then writing to researchers, then publishing. Two papers came out in 2021, one setting out thirteen hypotheses about which behavioural abilities appeared at which milestone across six hundred million years, the other introducing the five breakthroughs as a framework in its own right. Eventually he took a year off and turned all of it into A Brief History of Intelligence.
What he did was set three literatures next to each other.
Comparative psychology, which asks what various animals can actually do.
Evolutionary neuroscience, which asks in what order the bits of the brain showed up.
And AI, which acts as an unsentimental referee, because if you say you understand a mental process you ought to be able to build a shabby version and watch it run.
The book is about the 5 evolutionary breakthroughs that made us “intelligent”: steering, reinforcing, simulating, mentalising, speaking, and they developed in that order.
Steering
Why have a brain at all? Before there were brains there were nerve nets. Scattered reflexes, no centre, something along the lines of a coral polyp, which sits still and waits for food to drift into it. That arrangement worked for a very long time. Then, around six hundred million years ago, some animals started moving about, and one of them turned up with a small knot of neurons at the front end.
The stand-in for that animal is the nematode. Three hundred and two neurons, which is the entire nervous system, every connection of it mapped. No eyes. A few sensory cells for smell, light and pressure. Drop one into a dish with food in it and it will find the food nearly every time, with no picture of the world whatsoever. Keep going while the good smell strengthens, turn when it weakens, and let physics do the rest, since chemicals plume outward from a source and the concentration climbs as you close in.
They had a simple experiment to show how the brain began. Lay a strip of copper across the middle of the dish, copper being mildly toxic to nematodes and something they’d normally steer around, and put the food on the far side. The worm arrives at the barrier and weighs it up. Whether it crosses depends on how strong the food smell is and on how hungry it is. Hungrier worm, more willing. Stronger smell, more willing.
That’s the thing a nerve net can’t do. Somewhere the neuron excited by the food has to meet the neuron put off by the copper and be weighed against it on a single scale, so that one animal can produce one decision, and in the worm we’ve actually mapped the wiring that does it. Bennett’s understanding is that the first brains were built to make trade-offs. Everything else came later, and on top.
Two more details from the worm, both still running in you. Its dopamine neurons stick out of the animal’s body and detect food nearby, and when they fire the worm slows down and searches the local area, on the reasonable assumption that food comes in patches. Its serotonin neurons sit in the throat and detect food actually going down, and they produce rest. Dopamine for what’s out there, serotonin for what’s arrived. That got sorted out before anything had eyes and it is still, more or less, the arrangement that got you out of bed this morning.
Rodney Brooks, who invented the Roomba, put nearly the same algorithm into the first domestic robot anybody actually bought. Bump, turn, carry on, and when you find dirt, stop and search locally, because dirt also comes in patches. He got there by refusing to start at the complicated end.
Simulating
For a long time we give the neocortex credit for recognising the world. It’s the newest structure, the folded sheet you picture when you picture a brain.
Bennett then went and read the fish literature, which hardly anybody reads, because research money follows human disease and human disease means rats.
Turns out you can train a fish to squirt water at one particular human face for a reward. Rotate the photograph and it goes to the same face. Show it a frog and rotate that, and it picks the frog out first time. It’ll remember the trick a year later. Convolutional neural networks need absurd quantities of data to get near this and in some respects still don’t manage it.
A fish has no neocortex. It has a three-layered structure called the pallium, the ancestor of our hippocampus, our olfactory cortex, part of the amygdala.
So the question is if the job we’ve been crediting to the neocortex was already being done, and done well, by structures that predate it by a few hundred million years, then recognition can’t be the reason it evolved.
Bennett’s answer is simulation. Rendering a state of the world that isn’t the current one.
Location cells in a rat’s hippocampus fire according to where the rat is in a maze, and mostly they just track its actual position. David Redish recorded them at the moment a rat pauses at a junction and looks back and forth, and the cells stop reporting where the animal is and start running down the corridors it hasn’t taken. You can sit there and watch a rat imagine.
The same faculty turns up as counterfactual learning. Teach a chimpanzee rock-paper-scissors. It plays paper, you play scissors, it loses. Under plain trial and error the animal should now be equally more likely to play rock or scissors, the two moves that didn’t just lose. What it actually does is lean towards rock, the move that would have won, which only makes sense if something in its head reran the round with a different choice in it.
Reinforcement Learning
Edward Thorndike, early twentieth century, wanted to study children, wasn’t allowed to, and studied cats in puzzle boxes instead. He was looking for insight and for imitation. He found neither. What he got was a slow steady curve, animals trying things more or less at random and drifting towards whatever had worked. Reinforce recent behaviour when something good happens.
Marvin Minsky (the father of AI) built a machine on that principle in 1951, a maze-running rat made of vacuum tubes and clutches that strengthened whatever it had just done whenever it succeeded. It ran. It also went as far as that idea goes, which isn’t far. The problem is temporal credit assignment, and chess is a great example.
You win on move sixty, but the move that won it was probably move twenty-two. Reinforce the last few and you're rewarding whatever happened to be standing nearby at the end. Reinforce all of them and the signal spreads so thin across so many moves that you'd need more games than anyone will ever play.
Richard Sutton’s fix, was that the training signal isn’t the reward at all. It’s the change in expected future reward. Playing chess you’ve got one process picking moves and another one continuously estimating how well you stand, and what does the teaching is the moment the estimate jumps. That lift you get when your position suddenly improves is the mechanism, felt from the inside.
Turns out dopamine wasn’t tracking reward at all, it was tracking the revision to the estimate.
This is Wall Street’s version of “buy the rumour, sell the event”.
It also accounts for something I’d noticed for years: You chase something for three years and the day it lands is oddly flat. The lift came much earlier, in the week you first thought you’d probably get it, and by the time the thing itself turns up your estimate has long since moved and there’s nothing left for the circuit to pay you with.
Why The Order Matters
The dependencies are what make Bennett’s five steps more than just a list.
Temporal difference learning or learning to predict, where you train each prediction against your own next prediction instead of waiting for the actual outcome can’t happen without valence, because expected future reward needs some grounded sense of good and bad to be expected of.
Simulation can’t happen without reinforcement, since imagining a future is worthless unless you’ve got machinery that changes behaviour on the strength of an imagined outcome. Mentalising, which is modelling your own simulation, needs a simulation to model. And language needs all of it, because language moves the contents of an inner model around and is pointless if there isn’t one.
Early vertebrates learn from their own actual actions. Mammals add their own imagined actions, which is an enormous upgrade, because now you can be wrong in your head instead of in the world. Primates add other animals’ actual actions, by working out what they were trying to do.
The AI parallel doesn’t work because copying a human driver’s inputs directly produces a bad driver, because you faithfully reproduce all the little corrections and learn nothing about the intent, and what works instead is inferring the objective and training against that.
Four sources of learning, then. Own actions, own imagined actions, other people’s actions, other people’s imagined actions. Each one harder to reach than the last and each one a large multiplier on what a single life can hold.
Bennett comes down firmly with the majority view that language is an add-on. Thinking is the mammalian business of rendering a world that isn’t there, and language is the layer bolted on afterwards that lets two of those renderers compare notes.
What Bennett and a16z Agree On
Ask Bennett where he’d put his own research time now, he says continual learning.
Tell me something new and I know it, and I haven’t lost anything I already knew. LLMs don’t work like this. You don’t let the weights update continuously because you know you risk catastrophic forgetting. Fine-tune on a small set and it overfits to that set, loses generalisation, and drops capabilities it used to have.
His candidate mechanisms are interesting. Weights updated only under surprise, gated by neuromodulators that get released when a prediction fails. Possibly some offline consolidation phase, which he glosses, drily, as AI agents needing to sleep. He distinguishes all this from the more popular route of bolting persistent memory onto language models, which he describes as hacking the experience of continual learning onto the existing paradigm, while allowing that it might well produce something that feels like the real thing without being it.
In current AI models, “understanding” can sit in 3 different places. There are the weights, where anything resembling understanding lives, and which nobody dares touch for the reasons above. There’s the context window, holding whatever you’ve told it in this conversation, wiped when the conversation closes. And there’s a retrieval store off to the side, which persists, and which the model has to be handed again every single time because it has no idea the thing exists.
None of the three is what happens when you tell a person something. Your weights change, cheaply, without collateral damage. The nearest analogue in the machine is the context window, which is exactly why the illusion holds, because within a session it looks a great deal like learning. Then the session ends and none of it happened.
There’s an economic tail to this since the workaround has been to make the context ever larger, a learning problem has quietly become a memory bandwidth problem. Every token in context has to be held and read again on each pass, and it’s held in high-bandwidth memory sitting next to the accelerator, which is a large part of why HBM went from a commodity part to one of the tightest links in the entire chain. An extraordinary quantity of silicon is now dedicated to telling models things we have already told them.
Tell a transformer that one plus one equals four. Hand it a fine-tuning set built on that premise. It’ll learn it. It has no choice in the matter. You’re back-propagating through, and whether any of this ties with everything it already holds is not a question the mechanism knows how to ask.
Now tell a person. What happens is refusal. They start groping around for how it could conceivably be true, some other notation, a joke, a different base, and until they find a way to fit it into what they’ve already got, they throw it out.
This is why continual learning is key for our current model of AI to succeed. Even Andreessen Horowitz published a piece in April calling continual learning some of the most important work going on in AI, an idea they trace to McCloskey and Cohen in 1989. Isn’t continual learning, recursive self-improvement by another name?
From the piece:
“The thing that happened with AGI and pre-training is that in some sense they overshot the target… A human being is not an AGI. Yes, there is definitely a foundation of skills, but a human being lacks a huge amount of knowledge. Instead, we rely on continual learning. If I produce a super intelligent 15-year-old, they don’t know very much at all. A great student, very eager. You can say, ‘Go and be a programmer. Go and be a doctor.’ The deployment itself will involve some kind of a learning, trial-and-error period. It’s a process, not dropping the finished thing.”
— Ilya Sutskever
Why We Love Our IKEA Furniture
The fourth breakthrough is where the primates arrive, and where the book stops being about machines.
Compare a chimpanzee’s brain with a rat’s and there’s little new. Bigger, certainly. Structurally, two additions: a region of frontal cortex called granular prefrontal cortex, and a couple of new areas of posterior sensory cortex. Turns out that the granular prefrontal cortex gets no direct sensory input at all. Its information comes from the older frontal areas. A model of the model.
Your frontal cortex holds a model of you, assembled by watching what you do, and from your behaviour it infers your intent, and then it pushes you towards satisfying the intent it inferred. So the self gets built out of the behaviour rather than the other way round, which took me a while to accept.
The cleanest demonstration is an old cognitive dissonance study. Two groups watch the identical video about a group they’re about to join. One lot first has to read innocuous words aloud. The other has to read genuinely embarrassing words aloud. Same group, same information, different price of admission.
Afterwards the people who suffered more report liking the group more.
Which is why telling yourself what to want almost never works while doing the thing usually does. The model gets built from what you were observed doing, and stated intentions aren’t behaviour. It’s also why the expensive commitment ends up feeling like conviction rather than a performance of conviction. You paid for it, so something in you concludes it must have been worth paying for.
Maybe that’s why we love our IKEA furniture so much!
Going Back Down
Bennett doesn’t think the human brain is worth copying wholesale, and his reasons aren’t technical. Primates are political animals. We evolved obsessed with rank, watching who’s climbing, measuring ourselves against the troop all day long. That sits in the newest layer and it produces what he calls the worst versions of us. He’d rather nobody built it into a machine.
It’s a strange note for his story to end on, because each layer of intelligence arrived with its own problems. Simulation is what lets you plan a week ahead and it’s also what has you awake at four in the morning replaying a conversation from Tuesday. The self-model hands you an identity, and then you spend a good part of your life defending it from people who were never attacking. Most of what humans have invented for relief looks, from here, like a way of climbing back down the intelligence stack.
For example, what meditation does is drop you out of the primate layer, out of the noise about who you are and where you stand, and back to rendering the present and nothing else. Tolle would be proud.
One technique I have been thinking about since I read the book is around the self-model. If the part of you that reports your convictions is assembled from what you’ve already done, maybe it can’t be relied on to tell you when you’re wrong, because by then it has a stake in the answer. I will try to keep a page for every investment position of any size, written before I buy, saying what would have to happen for me to change my mind. I will read it when I’m irritated, which I was a lot last week.
Back to Bennett, turns out all five of his anchors of intelligence have been running while you read this.
Something inside of you sorted the page into worth-it and not, early, before you’d have called it a decision. Something else has been revising that estimate line by line ever since, which is why you’re still here, or why you left a while ago. You’ve been rendering worms and fish and a rat pausing at a junction, none of which are in the room with you. You’ve been modelling me at the same time, what he’s driving at, whether he’s overreaching, what sort of person writes this sort of thing. And all of it was set going by marks on a page, six hundred million years of machinery running on somebody else’s imagining.
Sources
Max S. Bennett, A Brief History of Intelligence: Evolution, AI, and the Five Breakthroughs That Made Our Brains (Mariner Books, 2023). abriefhistoryofintelligence.com
Max S. Bennett, “What Behavioral Abilities Emerged at Key Milestones in Human Brain Evolution? 13 Hypotheses on the 600-Million-Year Phylogenetic History of Human Intelligence”, Frontiers in Psychology 12:685853 (2021). doi.org/10.3389/fpsyg.2021.685853
Max S. Bennett, “Five Breakthroughs: A First Approximation of Brain Evolution From Early Bilaterians to Humans”, Frontiers in Neuroanatomy 15:693346 (2021) — the framework itself, peer-reviewed two years before the book. doi.org/10.3389/fnana.2021.693346
Max S. Bennett, Thomas P. Zollo and Richard Zemel, “Tell Me What To Learn: Generalizing Neural Memory to be Controllable in Natural Language”, arXiv:2602.23201 (February 2026, revised March 2026). arxiv.org/abs/2602.23201. Implementation at github.com/maxbennett/Generalized-Neural-Memory.
Bennett’s publication list and current affiliation: Google Scholar.
Andreessen Horowitz, “Why We Need Continual Learning” (April 2026). a16z.com/why-we-need-continual-learning
Interviews drawn on for Bennett’s own framing: Brain Inspired with Paul Middlebrooks, episode 181; The Cognitive Revolution with Nathan Labenz; and The Creative Process with Mia Funk, which is where the remarks on meditation, awareness and machine sentience appear.








