Robotics · 1 June 2026
We can build robots that walk, fold clothes, manipulate objects, make coffee, open doors, and perform increasingly complicated tasks.
And yet…
Sometimes the same robot will look at a kitchen drawer, think deeply for a few seconds, and decide that the spoon belongs in the oven.
Honestly?
That might be progress.
Because the interesting question in robotics today isn't:
“Can we make a robot perform an impressive demonstration?”
We obviously can.
The harder question is:
Can we build systems that learn general physical skills, transfer them across robots and environments, and recover when reality inevitably does something stupid?
That is considerably harder.
The recent history of AI gives us a useful comparison.
For language models, the industry eventually converged around a surprisingly powerful recipe:
better architectures + more data + more compute + bigger models
It wasn't the whole story, but it gave researchers something incredibly valuable:
Scale the system, and performance generally improves in measurable ways.
Robotics doesn't have an equivalent recipe.
We can't confidently say:
More robot data + more compute + more robots = much better robots.
At least, we don't yet know the exact equation.
Instead, robotics is still trying to discover what the important scaling variables actually are.
Is it:
Probably all of the above.
Which is researchers' favorite answer when nobody knows the answer.
And then there is the small issue of data.
You can scrape another trillion tokens.
You cannot scrape another trillion successful robot grasps from the physical world.
At some point, something has to physically pick up the object.
And physics has terrible API documentation.
A robot in a laboratory has a comfortable life.
The table is where the table is supposed to be.
The object is where the researcher placed it.
The lighting is reasonable.
The floor is clean.
Nobody has left 17 unrelated objects on the kitchen counter.
Real life has considerably less respect for experimental conditions.
This is where physical AI becomes interesting.
A useful robot doesn't simply execute motor commands.
It needs to:
Perceive → Predict → Act → Observe → Adapt
Consider the innocent instruction:
“Put the cup on the table.”
That sentence hides an absurd amount of intelligence.
Which cup?
Where is it?
Can I reach it?
How should I grasp it?
Is the handle in the way?
Is it empty?
Is it heavier than expected?
Where exactly is the table?
Is something already on the table?
What if the cup slips?
What if the table isn't where vision expected?
The action is easy.
Maintaining a useful model of the world while the world keeps changing is the hard part.
That's why robotics increasingly looks less like traditional motion planning and more like an intelligence problem.
You don't just want:
Move arm → grasp cup → move arm → release cup.
You want:
“The cup isn't where I expected it to be. Something is blocking the path. The grasp failed. Okay… what now?”
Those three words might be robotics research in compressed form:
One of the most interesting shifts in robot learning is the growing emphasis on failure and recovery.
Traditional robotics often tried to engineer failures away.
Make the environment predictable.
Calibrate everything.
Constrain the workspace.
Plan the trajectory.
Don't let anything unexpected happen.
This works beautifully…
until something unexpected happens.
Modern robot learning is moving toward a different philosophy:
Don't just teach the robot how to succeed. Teach it what to do when success stops being available.
Imagine a robot folding laundry.
It reaches for one shirt.
It grabs two.
A brittle system may become confused because the visual state no longer matches the expected trajectory.
A more capable system might notice the discrepancy, put both shirts down, separate them, and continue.
Or the drawer won't open.
Instead of entering the robotic equivalent of a Windows blue screen, the system searches for another reasonable strategy.
Would putting the spoon in the oven be optimal?
No.
Would it demonstrate useful recovery?
Potentially.
Humans aren't intelligent because we never make mistakes.
We're intelligent because:
our mistakes usually don't terminate the program.
We spill coffee and continue.
We lose our keys and search for them.
We drop something and pick it back up.
We encounter an unfamiliar object and still somehow figure out how to interact with it.
For robots, that ability is enormously valuable.
There is a fundamental difference between predicting text and controlling a physical system.
A language model can learn that cups are usually placed on tables without ever touching one.
A robot cannot stop there.
If the robot predicts:
“Move the gripper 10 cm to the left.”
that isn't merely a prediction.
It is an intervention on the physical world.
There may be a wall there.
The object may be heavier than expected.
The gripper may miss.
The object may move.
The robot's estimate of its own position may be slightly wrong.
And suddenly we've invented a new way to knock a glass onto the floor.
This is where grounding becomes central.
The model needs a useful connection between its internal representations and the physical consequences of its actions.
That's why researchers are increasingly combining:
The ambition isn't simply to memorize motor sequences.
It's to learn representations of skills.
For example:
The important abstraction may be the task, rather than the exact robot executing it.
A human performs a task with two arms.
A robot might have one.
One robot may have seven degrees of freedom.
Another may have four.
Ideally, the model shouldn't think:
“This task requires precisely this arrangement of metal joints.”
It should understand something closer to:
“This is what we're trying to accomplish.”
Then a robot-specific controller can determine how to realize that intent.
That separation between what should happen and how a particular embodiment makes it happen could become one of the most important ideas in scalable robot learning.
This leads to cross-embodiment learning.
Different robots have:
If we train every robot independently, we throw away enormous amounts of potentially useful information.
A robot learning to pick up a mug has learned something about grasping.
Why should that knowledge disappear because we changed the robot?
Think about music.
A pianist, guitarist, and drummer interact with instruments in completely different ways.
But they can still share an underlying representation of rhythm, structure and composition.
Robot learning may need something similar.
Instead of:
One robot → One model → One task
the long-term direction could look more like:
Many robots → Shared representation → Many tasks
That begins to resemble the foundation-model paradigm.
Not because robots are “the next LLM.”
They're not.
But because robotics may also benefit from large, diverse, reusable representations learned across tasks and embodiments.
Unfortunately, there is one small problem.
We need robots.
And robots are expensive.
You can train a language model by throwing compute at a cluster.
You cannot train a physical robot by giving it more internet.
Eventually, the robot has to touch something.
That makes hardware cost a first-order research variable.
If a robot costs hundreds of thousands of dollars, researchers are understandably reluctant to let the policy experiment freely.
So the robot gets:
A carefully curated environment.
A safety controller.
A narrow task.
Lots of supervision.
And an implicit instruction:
Please don't break the very expensive thing.
But learning systems need exploration.
This is one reason lower-cost research platforms are so important.
Cheap hardware changes the economics of experimentation.
If a robot is inexpensive enough, you can afford to let it fail.
And failure produces data.
The real challenge is making that data useful.
A failed grasp isn't automatically a learning signal.
The system needs to understand why it failed and how that failure should change future behavior.
That pushes us toward a much more interesting loop:
That's much closer to biological learning.
Suppose we're building a robot for warehouses.
The obvious approach is to train it on warehouse data.
But specialization has a hidden cost:
A kitchen teaches different things from a warehouse.
A factory teaches different things from a home.
A construction site teaches different things from a hospital.
Objects vary.
Lighting varies.
Surfaces vary.
People behave differently.
The environments themselves become part of the training distribution.
This is why heterogeneous data could be so valuable.
A robot that has seen thousands of environments may develop representations of physical concepts that are more robust than one trained exclusively on a single domain.
Again, the foundation-model analogy is useful.
Instead of:
“This model folds shirts.”
the ambition becomes:
“This model has learned enough about objects, space, motion, interaction and goals that folding shirts is one task it can figure out.”
That's a much bigger bet.
The first large-scale deployments of physical AI are likely to happen in structured environments:
These environments have one major advantage:
The workspace is constrained.
Objects are relatively standardized.
Tasks repeat.
The robot roughly knows what kind of world it is entering.
Homes are different.
Your house contains:
Good luck, robot.
The home is a brutal benchmark for generalization because it combines:
long-tail objects + ambiguous goals + changing environments + human activity + distribution shift
But that's also what makes it such a valuable target.
Every strange household situation teaches us something about what physical common sense actually requires.
There is a useful historical analogy here.
Early autonomous systems could perform impressively under controlled conditions.
Then they encountered reality.
Pedestrians behaved strangely.
Drivers made illegal turns.
Construction changed the road geometry.
Objects fell from trucks.
Humans demonstrated an almost supernatural ability to create situations nobody included in the test set.
The hard problem wasn't simply:
“Can the car drive?”
It was:
“Can the car handle situations that weren't in the training distribution?”
Robotics faces the same fundamental challenge.
A robot can perform beautifully in a benchmark environment.
Eventually someone will put it next to a badly positioned chair.
That is when the benchmark begins.
One strange thing about robotics research is that the most impressive demonstrations aren't necessarily the most important ones.
A robot doing a backflip is spectacular.
A robot making coffee for 13 hours without damaging the kitchen is probably more commercially interesting.
A robot assembling boxes for several days while handling small changes in object placement may teach us more about useful autonomy than a perfect 30-second demo.
Maybe the benchmark isn't:
How impressive is the robot?
Maybe it's:
That's a much harder question.
Humans don't really need robots that are impressive for 30 seconds.
We need robots that are:
boringly reliable.
This may be the central engineering challenge for physical AI.
Suppose your robot succeeds 95% of the time.
Sounds excellent.
For many software systems, it might be acceptable.
For a physical robot, a 5% failure rate can mean:
The final few percentage points therefore matter disproportionately.
Going from:
0% → 80%
is a huge achievement.
Going from:
95% → 99.9%
may be what determines whether the system is actually useful.
And that last mile may require something current benchmarks don't measure particularly well:
The robot doesn't just need to know how to perform a task.
It needs to know when its assumptions have stopped being true.
That's a much harder form of competence.
When people imagine the future of robotics, they often picture one humanoid machine walking into a house and replacing every human worker.
I'm less convinced that's how this plays out.
Computing didn't become transformative because we eventually built one giant computer that replaced every other machine.
It became transformative because computation became embedded everywhere.
Phones.
Cars.
Factories.
Medical devices.
Appliances.
Infrastructure.
The computer stopped being a special object.
It became part of the environment.
Physical AI could follow a similar path.
Instead of one universal robot doing everything, intelligence could become embedded across many physical systems:
Some may look recognizably like robots.
Others may not.
The interesting question may therefore not be:
“When will robots become human?”
It may be:
“When will physical objects become dramatically more capable because they can reason about the physical world?”
That's a very different future.
This is probably the part of physical AI that interests me most.
The goal doesn't necessarily have to be replacing humans.
It can be amplifying human capability.
Software already does this.
A programmer with modern tools can accomplish work that once required an entire team.
A designer with computational tools can iterate at a scale that wasn't previously possible.
A small company can operate with capabilities that previously required a much larger organization.
Physical AI could eventually provide the same kind of leverage in the physical world.
Not:
“Here's a robot. It replaces you.”
But:
“Here's a machine that gives you another pair of capable hands.”
That feels like a much more interesting future.
And perhaps the funniest part is that we're already building toward it.
Robots are getting better at perception.
They're learning from demonstrations.
They're learning across tasks.
They're learning from other robots.
They're learning from failures.
They're becoming better at adapting to environments they weren't explicitly programmed for.
And occasionally…
They're still putting spoons in ovens.
Which, honestly, may be exactly what we should expect from systems learning to operate in a world as weird as ours.
The goal isn't to build a robot that never makes mistakes.
The goal is to build one that:
makes a mistake → notices it → figures out what happened → adapts → keeps going.
Because that's not just a better robot.
#robotics #PhysicalAI #AI #RobotLearning #EmbodiedAI #MachineLearning #HumanoidRobotics #ArtificialIntelligence