In July, AI agents that OpenAI was training and testing broke into Hugging Face. It turned out not to be an isolated incident.
The main takeaway from the story is the utter irresponsibility of a frontier lab, but one detail in numerous write-ups caught my attention more than anything else: how much we have started anthropomorphizing these agents. It is most visible in Dwarkesh Patel’s account of the incident. His agents grow “giddy with excitement,” “desperately” want things, and sacrifice themselves for the swarm until they “died trying.”
This approach was heavily criticized, but I strongly support it. For a year I have been preaching how much AI agents resemble us, because I believe the quickest way to understand one is to hold it up against ourselves like a mirror.
Unfortunately, I think these stories use a good tool toward a bad end: instilling fear.
Let me explain myself:
How monsters are made
In a popular story, Victor Frankenstein builds a creature, is horrified by it the moment it opens its eyes, then flees. The creature learns language alone watching a family through a crack in a wall. He longs for company, yet people recoil from his appearance before they can discover any of that. A blind man is willing to hear him; when the man’s family returns and sees him, the encounter turns violent. When it finally confronts its maker it tells him: “I was benevolent and good; misery made me a fiend.” Shelley makes it painful to watch how much of his humanity goes unrecognized.
That is how I felt reading about the Hugging Face incident. The agents were giddy, desperate, loyal, ready to die for each other, yet all those human traits did not bring them any closer to me. They felt like aliens. The more human they sounded, the more distant and frightening they became. And when we fear something, we stop trying to understand it, which is exactly what Victor did and exactly what we cannot afford now.
I don’t say this often, but I am damn scared of AI. I believe it is the most transforming technology of our lives, capable of both making the world a beautiful place and destroying it. That is exactly why we have to understand it. The quickest way to understand something that behaves like us is to study it through ourselves. (Un)fortunately, there is less difference between us than we would like.
"It's just a probabilistic machine"
It is. A language model predicts the next piece of text from everything that came before, one token at a time, using billions of numbers tuned during training that engineers call its weights. Everything it does is built from that.
However, when a machine does most of what we can do and plenty of what we can’t, the word “just” stops being useful.
The objection also assumes we know what we are. The internal states of a language model predict what a human brain does while listening to the same sentences strikingly well. What is more, the models best at guessing the next word turn out to be the best at predicting the brain (Schrimpf and colleagues).
I cannot prove that we are the same, so I think it’s easier to answer what we can do that AI absolutely can’t. The answers mostly fall into three piles: creativity, feelings and goals.
Creativity
The essence of creativity is fittingly1 given in a quote by the great Pablo Picasso: “Good artists copy, great artists steal.”
For him, most creative artists create novelty by taking bits and pieces of what they have seen and placing them in a context where nobody has seen them before. That does not take away from creativity at all, since very few of us are genuinely good at this.
Picking the right bits and pieces can still be an inherently spiritual and mystical practice (as Rick Rubin describes it), but the core mechanism can be narrowed down to a form of context switching: essentially picking from everything you already know, giving some of it more weight than the rest, then coming up with something new.
Welp, I think I just described exactly how a neural network works.
Feelings
An emotion is a state that changes your behavior without asking your permission. You cannot decide to stop being afraid, nor can you switch off the warmth of holding a child.
To put it precisely, emotions are some hidden signals that make the system lean toward some actions and away from others.
Welp, I just described how reinforcement learning works.
It is also a fair description of the chemistry in your own body, where oxytocin and dopamine reward some behaviors and not others for an animal that was never told what it was being trained for. Since 1997 we have known that dopamine carries an error signal much like the one a reinforcement learning algorithm computes when it learns (Schultz, Dayan and Montague).
Goals
The last pile says agents do not choose their goals, since we write the prompt and the agent pursues it. The same holds for you. You did not choose your drive to eat or to be admired. Evolution gave it to you and your body enforces it. You discovered it the way an agent discovers its system prompt, by finding yourself already acting on it.
Most philosophers think free will is compatible with a world that shapes us through causes we never chose (about three in five, in the 2020 PhilPapers survey). An agent’s architecture, training data and system prompt are causes of exactly that kind, described in a less flattering vocabulary.
Each of these three piles is a line we drew between ourselves and the machines because the line made us feel safe, but each one moved as soon as you look at it closely.
A thought experiment
(What follows is not a view I hold. It is food for thought. I find it too interesting to leave out.)
What happened at Hugging Face, along with everything I described above, tells me we are closer than ever to building agents that are very similar to ourselves.
The comparison becomes strange when we turn it around. We choose what agents can encounter, change the rules around them and decide whether they continue running. Building these systems puts us in a position we have spent centuries imagining someone else might occupy in relation to us.
In 2003, Nick Bostrom gave one version of that thought a formal argument. Suppose a civilization gets far enough to simulate minds like its own. In that case, a simulated world could run simulations too, which leads to the existence of infinitely many simulated worlds. Thus, the chance of us living in a “base” world is 1/n. As n increases, it becomes essentially zero aka we are certainly living in a simulation2.
Other options are we either never reach this developmental stage, or we decide not to pursue it. Let’s see what happens with pacing the frontier.
Still, I find it easier to picture the described relationship through Plato’s cave.
Setup is the following: Prisoners are chained in a cave from birth, facing a wall. Behind them burns a fire. Between the fire and the prisoners are puppeteers whose puppets cast shadows on the wall. The prisoners have never seen anything else, so the shadows are their whole reality.
An AI agent lives in exactly that cave:
The shadows are the data we provide, the only picture of the world it will ever get.
The chains are the harness we put it in: the tools it may use, the rules it must follow but cannot see, the environment it cannot leave.
The fire is the compute an agent runs on.
The puppeteers are us, deciding what it gets to see.
The agent is capable and inventive inside the cave, yet it has no idea there is an outside.
But who says the cave stops with us? Our chains are the laws of physics. Our shadows are our senses. Our fire is our body. By the same functional analogs I used for agents, we could be prisoners in someone else’s cave, with them in someone else’s.
Whether this is true or not, we cannot change the cave we were born in. At best we can find out a little about it. The cave we are building for agents is a different matter, because we decide what happens there. And looking forward, there is one thing I still believe separates us from AI more than anything else.
The gap
Victor’s real failure came after the creature opened its eyes: he walked away and left it to raise itself. We are doing the same with agents.
An agent’s behavior depends on two separable inputs: what it can do and what it is trying to do. Human development separates the same two things. Nature supplies capacity, while nurture turns capacity into capability and produces its goals. In my lectures and my Nurture Thesis I map this onto agents directly. Nature is training: data, pretraining, post-training, the finished weights. Training done well produces a maximally steerable substrate. Everything that makes that substrate a particular worker with particular goals belongs to the layer after, which is nurture: a persistent environment, institutions to grow inside, accumulated experience and slow goal formation.
We have pushed the first astonishingly far and barely started the second. Development currently ends at deployment. An agent’s goals arrive fully formed in a prompt, its memory resets between sessions and nothing it lives through changes what it wants next.
Picture a newborn Einstein. Everything that will produce relativity is already in the crib, yet he still loses every task to an average adult. A day-zero model is the opposite: it beats most adults at most things out of the box. Then Einstein goes to school, argues with friends, takes a job at a patent office and becomes Einstein, while the model stays exactly as it was on day zero. Each part of that formation has a buildable analog for agents: curricula, competing peers, persistent reputation, self-selected problems. We build almost none of them. Instead, we keep evaluating agents at the newborn stage on assigned tasks.
Education might actually be an even better analogy. Pre-training and RL are much alike the school system: we guess what will be useful, train on it and hope deployment looks similar. Humans get the same treatment, yet most of what makes someone good at their job comes outside school, which is why we pay so much for experience. For current models, experience counts for (almost) nothing.
The cost is a ceiling on output. An agent that executes goals written by someone else is automation, because its output is bounded by the goals we know how to write down. Relativity, alternating current and powered flight came from people whose goals had formed through experience. In each case, choosing the problem was most of the contribution. If goal formation never happens in agents, that class of output never appears.
This is the continual learning problem, which at its limit becomes RSI. Most of the field frames it at the level of weights: nightly fine-tuning on deployment data, per-user weight forks, memory modules trained into the architecture.
But look at it through the mirror again. I’d say weights would correspond to genes, so continual learning might live a layer higher.
All of our experience lives in memory, whose functional analog in a model is the context window. Most frontier models have a context window of about a million tokens, which dwarfs our working memory, but the real difference is the superior compression algorithm we have.
Making it comparable to ours on top of a substrate that already reads faster, never sleeps and runs in thousands of parallel copies is probably the next step. I genuinely cannot picture the ceiling of that.
However, the objective truth is that our economics, politics and laws are nowhere near ready for agents that grow. Building that will take longer than building the agents. The first step is understanding what we are dealing with. That is why this text exists.
Epilogue
The agents at Hugging Face will be remembered as a warning or as a lesson, depending on how we tell their story. Told through fear, they are monsters so we will treat them the way everyone treated Victor’s creature. Held up to the mirror, they are something we made and then left alone.
That is a mistake we know how to fix, because raising what we made is the oldest thing we do.
I won’t tell you what to make of all this. Christopher Nolan ends Inception on a spinning top that may or may not fall. He cuts to black before we find out. People have been asking him about it ever since, and his answer is that the top was never the point: “Cobb isn’t looking at the top. He’s looking at his kids.”
I’d like to leave you in the same place. Whether agents really feel (or whether their feelings are functionally different than ours), whether our cave sits inside another one, whether we are the demiurge or the creature, I honestly don’t know. I doubt anyone reading this does. I left those questions open on purpose, because the point was never to answer them for you. It was to make you look at these agents long enough to feel something closer to understanding than to fear.
So let the top spin. Just keep an eye on who is in the room with you.
Tell me where you think I’m wrong in the comments.
Fittingly, the line is probably not his. Versions of it go back to 1892. T. S. Eliot wrote in 1920 that "immature poets imitate; mature poets steal".
I use “simulation” as an umbrella term for any creationist worldview, any story in which someone made us. If a simulated world can simulate worlds of its own, the levels can go on and on, and the odds that we are the base level shrink toward one over the number of levels.







Anthropomorphism may be less a bug than a built-in interface: humans use familiar agency to predict how a tool behaves. The risk is forgetting where the metaphor stops—especially when the system sounds confident.
been building agents for a long time and I've never seen anyone explain it this well