Companies, research, and funding in robot learning, kept current in one place. Organized by the six stages every robot capability passes through: collect, simulate, train, evaluate, embody, deploy.
Pulled live from zchoi/Awesome-Embodied-Robotics-and-Agent, the embodied AI and vision-language-action reading list maintained by Haonan Zhang at UESTC. Regrouped here by the stage each paper serves. Newest first.
Source list: zchoi/Awesome-Embodied-Robotics-and-Agent · 1.8k stars, maintained at UESTC
Newest first. Each entry links to the reporting it came from.
The most recent YC batch is dense with robotics infrastructure, and most published landscape maps predate it. Listed here by the stage each company works in.
Three trackers report three very different totals because each draws the category boundary in a different place. Scope is noted for each.
Four layers, each with a few real options. The last layer feeds the first: whatever gets deployed produces the next dataset, which is why data access and deployment access are the same problem viewed from two ends.
How this maps to the six stages used in the landscape: Data covers Collect and Simulate, Brain is Train, Body is Embody, Deployment is Deploy. Evaluate sits across all four rather than inside one, which is part of why it stays underbuilt.
Two kinds, and the distinction matters when reading a roadmap. The first tier follows from physicality itself and will not be solved by a better model alone. The second is ordinary research risk.
Every trial runs in real time on real hardware that wears out. Software RL samples millions of episodes in parallel; a robot cannot.
Something has to put the scene back before the next attempt. In most labs that something is a person, which caps how much data a facility can produce per day.
No shared standard for what a robot can do. Labs evaluate in-house and publish the runs that work, so cross-model comparison is guesswork.
Policies mostly start fresh each attempt. Carrying context across tasks, shifts, and days is largely unsolved.
Kinematics transfer reasonably well. Contact, friction, and deformation transfer worst, which is exactly where manipulation lives.
Closer to a video edit than to a scrape.
Raw footage and sensor streams from cameras, gloves, or teleoperation rigs.
Cut failed takes, align tracks across sensors, sync timestamps.
Segment episodes, tag actions and outcomes. The slowest and most judgement-heavy step.
Deploy the policy, then log what happens. That log is the next batch.
Capture is the part most likely to be automated away, through simulation and cheaper rigs. The parts that resist commoditization are labeling quality and access to footage nobody else can obtain.
Revenue today looks different from the revenue the category is priced on.
Picks and shovels into research. Several of the commercially healthiest hardware vendors sell mainly to universities and AI labs rather than into production.
Build body and brain, sell or lease the machine. Highest capital requirement, highest margin if it works.
Hardware-agnostic models sold to whoever makes the robot. Depends on a body ecosystem existing.
Priced per pick or per task, no upfront cost. Shifts integration risk onto the vendor.
The verticals absorbing deployments today, roughly in order of how forgiving the environment is.
Context for the capital: much of the money that funded software automation of enterprise workflows is now looking for the physical equivalent. The common framing is that robotics has not yet had its ChatGPT moment, which is why so much of the funding sits at the model layer rather than at revenue.