I am always thinking about I, Robot, not because I think the robots are about to rise up and kill us, but because I think Asimov may have accidentally given us a way of seeing the problem that no longer works.
When I say robot, what I mean in my head is embodied AI. I don't mean a dumb machine on a manufacturing line that has been programmed to repeat the same movement ten thousand times. I mean intelligence with a body. Something that can perceive the world around it, make some model of what is happening, decide what to do next, and then physically act on that decision. Just this morning I mentioned embodied AI to a group I am working with at the Oklahoma AI Roundtable because I think this distinction is becoming increasingly important. A robot isn't necessarily just automation anymore. Increasingly, it is AI that can reach out and touch the world.
Most of us who grew up on science fiction have some version of the Three Laws rattling around in our heads, whether we remember the exact wording or not. A robot can't hurt a human. It has to obey humans unless doing so would hurt one. It has to protect itself unless that conflicts with the first two. The stories get interesting because the laws collide, or because a robot finds some interpretation of them nobody anticipated. But underneath all of it is this assumption that I think about all the time: somewhere inside the machine there are rules.
That isn't really what we're building anymore.
I don't mean that literally, of course. There are still enormous amounts of deterministic software in a modern humanoid robot, and there had better be. Motors have controllers, batteries have management systems, safety systems have thresholds, networks have protocols. But the thing increasingly deciding what the robot sees, what matters in that scene, what might happen next, and what action might accomplish the task isn't a giant decision tree somebody sat down and wrote. It's a learned model.
And this is where world models start to matter.
Fei-Fei Li has been talking for years about spatial intelligence and the idea that language alone is not enough for an intelligence that has to operate in the physical world. Her work at World Labs is built around that problem. An AI needs some internal representation of space, objects, relationships, and eventually dynamics. It needs to understand that the coffee cup is behind the laptop even when it can't see the whole thing, that the table continues to exist when it turns its head, and that if it pushes something toward the edge gravity is going to finish the job. World models are attempts to give machines some version of that internal understanding of the environment, not merely recognize what is in an image but predict how the world might change.
Yann LeCun has been attacking essentially the same problem from another direction. His work on JEPA-style world models is based on the idea that intelligent systems need to learn how the physical world behaves so they can predict before they act. Meta's V-JEPA 2, for example, was trained largely from video to learn representations of physical behavior and has been used for robot planning in environments and with objects the system wasn't specifically trained on. In other words, we're moving toward machines that don't simply recognize a scene and retrieve a programmed response. They develop enough of a model of the world to ask, in effect, "If I do this, what happens next?"
Then there is simulation. If you can build convincing enough virtual environments, you don't have to teach every robot everything by banging expensive hardware into walls for three years. You can let models experience enormous numbers of variations in simulated worlds, learn policies there, and then transfer some of that learning into physical machines. Figure is going even further by pretraining its humanoids on large amounts of human behavior data and then testing whether those learned behaviors transfer into places the robot has never seen before.
That is a very different idea of a robot from the one most of us grew up with.
And I think that difference matters a lot more than we're talking about.
We've spent the last few years getting accustomed to probabilistic software because most of the time it lives behind a screen. ChatGPT hallucinates something and you roll your eyes. An image model gives somebody six fingers and we laugh about it. An agent screws up a browser task and maybe somebody loses an afternoon fixing whatever it did. There are certainly cases where software causes real damage, and there always have been, but there is still this intuitive boundary in our heads between software doing something stupid and a machine physically doing something stupid in the room with us.
We're crossing that boundary now.
Tesla is pushing Optimus toward production. Boston Dynamics isn't just making Atlas videos that half the internet watches because the robot moves in ways that make some primitive part of our brain uncomfortable. Uncanny Valley, anyone? Atlas is becoming a real industrial product. Figure is training humanoids with neural systems that can generalize movements and tasks instead of requiring somebody to script every movement in advance. Schaeffler has plans for more than a thousand humanoids across its manufacturing network.
And then there is proprioception.
Figure's newer systems connect vision, touch, and proprioception directly into whole-body control. Proprioception is basically your body's subconscious sense of where it is in space. You don't have to look at your hand to know where your hand is, or consciously calculate how much force your leg needs when you stand up. Your brain and body just know. It is sometimes called our sixth sense, and we're now trying to give a version of that sense to machines.
Think about what that means for a second. We're not just teaching a robot to recognize a coffee cup. We're teaching it to understand where the cup is in relation to its own body, where its arm is without having to stare at it, how far it needs to reach, how much force to apply, and how its own movement changes the world around it. Vision tells it there is a cup. Proprioception helps it understand, in a very primitive sense, where it is in relation to the cup.
That starts to look a lot less like traditional automation and a lot more like embodiment.
I don't think any of that means we're living in Terminator. I actually think that comparison gets in the way.
Johnny Five is alive, though.
What interests me is that software is getting a body.
And once software has a body, the old question of "why did it do that?" gets much more interesting.
By Matthew Williamson · Written for crossinginto.ai · October 7, 2026