Apr 2026 – Present
Robotics foundation models are the strongest starting point we have for manipulation, but none of the open-weight ones arrive ready for your robot. Unless your setup matches something already in the pretraining mix, the model needs demonstrations from that exact embodiment: those cameras, that mounting geometry, that gripper. They come from a human teleoperating the arm one trajectory at a time, and the whole collection is repeated for every new task, gripper, or camera placement.
This project removes the human from that loop. Pi 0.5 is fine-tuned entirely on demonstrations generated in NVIDIA Isaac Sim, with cuRobo planning the motions, then deployed zero-shot on a real Franka, with no teleoperation and not one frame of data from the real setup. The task is to grasp and pick up: the policy is given an object’s name and has to lift it off the table. It succeeds 80% of the time on hardware it has never seen.
Grasping makes a clean test because it is decided by vision alone: the gripper has to reach the right position and orientation, and the object’s location is known only from the camera images. Doing it on real hardware establishes that the visual sim-to-real gap is closed, and the same recipe extends to more complex visually guided manipulation.
How many simulated demonstrations this takes, and how that compares with collecting on real hardware, has also been analysed.

The real cameras are calibrated first, giving the 6D pose of each one in the robot base frame. Those poses are reproduced in simulation so the viewpoints match.
15 objects, each episode paired with a randomized prompt naming the object to pick. cuRobo plans the motion and grasps come from object geometry and inertia, so nothing is teleoperated. Domain randomization covers light (intensity, temperature, position), object texture, background, and camera pose.
PhysX runs at 60 Hz and rendering at 30 Hz, both recorded at 30 Hz, converted to LeRobot v2.1 and pushed to Hugging Face. One NVIDIA RTX 6000 Ada Generation collects an episode every 2 minutes. Following simplifications have been done to make the simulation faster.

Pi 0.5 is fine-tuned with a fork of openpi, conditioned on robot state and three cameras, predicting joint angle deltas rather than absolute targets. The SigLIP vision encoder trains fully unfrozen, while the VLM backbone decoder is LoRA fine-tuned at rank 16 and the action head at rank 32. NVIDIA H100, 1.3 s per step.

Policy server on an NVIDIA RTX 6000 Ada Generation (48 GB); the laptop at the robot streams cameras and state over a WebSocket and receives action chunks, driving a Franka FER arm at 10 Hz.
Real-time chunking (paper) keeps the arm moving: the next inference fires while the current chunk is still executing and is sent back with the request, so the sampled chunk agrees with motion already committed. No stop at chunk boundaries.
Two client-side fixes, since these artifacts are fixed amplitudes in radians and their implied acceleration scales as 1/Δt²:
Joint logs are sampled at the driver’s native 1.4 kHz, not on the 10 Hz policy tick, which would alias away the band vibration lives in.
How many demonstrations does a task need? Sim-to-sim answers that first: train in simulation, test on placements the policy was never shown, and read off the episode count where it starts working. That is the baseline the sim-to-real requirement is measured against.
Two tasks set the scale. Picking up a 4 cm cylinder needs the gripper precise in x and y, since a cylinder looks the same from every direction. Picking up a cuboid adds yaw. Nothing else changes between them, so the gap between the two episode counts is what one extra degree of freedom costs.
Both are scored over a 50 by 15 cm patch of table, at placements the policy was never shown.
Position first. Fifty episodes spread over 5 sites fails: half the objects fall in the gaps between sites. Thirteen episodes over 13 sites, a quarter of the data, picks the object up everywhere it is tested.
What counts is the spacing between demonstrations, not how many there are. Thirteen sites over the 750 cm² patch is one demonstration per 58 cm². At that spacing the policy covers the gaps on its own; at one per 150 cm² it does not. And more data at the same spacing adds nothing: 65 episodes over those same 13 sites still scores 100%.
Yaw costs ten times more. The cuboid needs 130 episodes, the same 13 sites with 10 angles at each, and the angles have to differ. Ten episodes split over two angles fails 60% of the time on anything else, over four angles about a quarter of the time. Only a fresh angle every episode works everywhere.
So 13 episodes for x and y, 130 once yaw matters too. Every axis the gripper has to get right multiplies the data, and a task needing a full 6D pose adds roll and pitch on top of that.

Simulation said 130 episodes. The real robot needs more, and the way it fails at 130 shows why. It does not reach for the object and miss by a little. It moves off in a completely wrong direction, as if it cannot see the object at all.
At 390 episodes that stops. The robot reaches the object, and what is left is near misses and objects dropped on the way up. The score is 60%. At 650 episodes it reaches 90%. That is five times the simulation number.
By 650 the policy has seen enough variation in light, table, object colour and camera position to stop depending on any of them. That invariance is what carries it over to the real robot.
It is a property of the dataset, not of the task. The same 650 episodes split across 8 objects still gives 80%, because every episode of every task randomizes the same way. So the cost is paid once: any task, or group of tasks, that would have needed around 650 real demonstrations anyway needs no extra episodes in simulation.
Around 650 episodes of domain randomization is what it takes for the policy to stop depending on light, table, colour and camera position, and that invariance is what makes it work on the real robot. It is a one-time cost. It does not grow with the number of tasks, because every episode of every task is randomized the same way.
Everything past 650 is not a sim-to-real cost. It is what the tasks themselves need, and it would be needed on real hardware too. A mustard bottle has several valid grasp modes. A facewash bottle is tapered, so a top-down grasp slips off it. cuRobo returns different trajectories for the same grasp pose depending on how close the object sits to the robot base. Each of these adds trajectories the policy has to see, in simulation or on hardware.
So once a dataset is around 650 episodes or larger, simulation costs no extra episodes. Here that dataset is 15 objects and 10 000 episodes, reaching nearly 80% on the real robot, with recoveries and with grasps at positions and orientations that were never trained on. Objects that were not in the dataset work too when their geometry is close to one that was. None of it is teleoperated.
The same recipe applies to longer trajectories and multi-step tasks, as long as the gap is only visual. What changes is how much data is needed, not the method.
Contact-rich tasks need more. The visual part carries over, but forces decide success, so randomization has to extend to friction, mass and inertia, restitution and contact stiffness. The collision approximations also have to be revisited: a coarse convex decomposition is enough when contact only has to be plausible for a grasp, and not when contact is the task.