Simulation enables large-scale, low-cost robot data generation, but policies trained in simulation often fail to transfer to the real world due to the sim-to-real visual discrepancies. Existing approaches often rely on intermediate representations, which can discard rich semantic information or require additional perception modules at deployment. We address this visual sim-to-real gap with RoboRender, a framework that converts simulated trajectories into photorealistic RGB videos for policy learning. RoboRender trains a robot-oriented video generation model conditioned on simulated depth videos, language instructions, and robot RGB mask videos, preserving simulator geometry, robot motion, and action labels while synthesizing realistic textures, backgrounds, and distractors. The generated RGB videos are paired with simulator-provided states and actions to train policies for zero-shot real-world deployment. On robot video test sets, our video model outperforms depth-conditioned video generation baselines in generation quality. In real-world experiments across pick-and-place, articulated-object manipulation, and mobile manipulation tasks, policies trained on RoboRender-generated data achieve a 71% average success rate, outperforming raw simulation renderings and conventional visual domain randomization by approximately 7.1× and 3.6×, respectively. We further show that policy performance improves with more generated videos per simulation trajectory, increasing opening-task success by 65 percentage points. These results demonstrate that generative video rendering mitigates the visual sim-to-real gap for zero-shot policy transfer.
RoboRender adapts Wan2.1-T2V-1.3B into a robot-oriented renderer, fine-tuned on ~130k real robot manipulation samples (~70k from AgiBot-World-Alpha, ~60k from DROID) at 416×240 — see Training data for the corpus and how it was built. The model is trained on real video and only conditioned on simulation at generation time.
Rendering a whole simulation dataset takes many diffusion passes, so inference cost bounds how much data you can generate. Two optimizations bring it down:
Existing robot datasets are collected for policy learning, not video generation, so they lack the prompts, geometric conditions and embodiment signals this model needs. Our preprocessing pipeline turns existing demonstrations into training-ready samples: for each clip we take a temporally consistent 81-frame window across all camera views and derive four aligned conditions.
A vision-language model runs in two passes: a draft caption from the first frame, then a refinement that cross-verifies objects across all non-wrist views. Constraining the output to structured metadata cuts hallucination substantially versus single-pass captioning.
Video-Depth-Anything for all samples, plus FoundationStereo where stereo recordings exist, both normalized by inverse-depth mapping. Depth sources are randomly mixed per view during training, so the model does not inherit the bias of any single estimator.
SAM3 segments the arm and gripper across every view, isolating the agent's pose and motion from the rest of the scene so embodiment can be varied independently of the background.
The frame immediately before the 81-frame window. This is the temporal conditioning the model chains on at generation time to run rollouts past its fixed frame window.
Distribution of the 16 task types across the training corpus. Hover a slice or a row for its sample count.
| Source | Samples | Distinct scenes | Task categories |
|---|---|---|---|
| AgiBot-World-Alpha | 69,910 | 22,650 | 10 |
| DROID | 57,040 | 17,042 | 15 |
| Total | 126,950 | 39,692 | 16 |
Distractors are inserted by cropping the condition, not the scene. Temporally consistent rectangles are punched out of the simulation depth map — the black patches below — so those regions are no longer pinned to simulated geometry, and the model fills them with clutter from the prompt: stable over time, and clear of the geometry the task depends on. For DROID pick-and-place we crop 2–5 patches covering 20–50% of the external view and 1–2 in the wrist view.
We evaluate on 500 held-out clips from the real training corpus (DROID and AgiBot), under two inference-matched settings: first-chunk generation, which runs without previous-frame conditioning, and autoregressive continuation, which is conditioned on a ground-truth previous frame. The baselines are three depth-conditioned video generators — Wan-Fun-Control, Cosmos-Transfer2.5 and RoboTransfer. The clips below show a qualitative comparison against the first two.
Open the fridge
Open the drawer
| First-chunk | AR continuation | ||||
|---|---|---|---|---|---|
| Model | FID ↓ | FVD ↓ | PSNR ↑ | SSIM ↑ | LPIPS ↓ |
| Wan-Fun-Control | 62.5 | 267.8 | 20.30 | 0.738 | 0.175 |
| Cosmos-Transfer2.5 | 42.8 | 164.5 | 13.26 | 0.481 | 0.423 |
| RoboTransfer | 52.8 | 275.4 | 13.75 | 0.509 | 0.418 |
| RoboRender | 25.7 | 93.9 | 22.28 | 0.784 | 0.142 |
Per-frame PSNR and LPIPS along 15-second autoregressive rollouts, sampled every 3 seconds — quality holds up across the repeated steps needed to render a full trajectory.
Every policy here is trained purely in simulation and deployed zero-shot — no real-world fine-tuning, and no real data beyond what trained the renderer. Each is a fine-tuned π0.5 predicting 32-step action chunks with the first 16 executed, trained for 40k steps at batch size 72. Evaluation runs in a natural office environment, where object size, geometry, camera viewpoint and background all differ from simulation, across eight tasks and three embodiments.
The policies are trained on simulation trajectories, which are separate from the real corpus that trained the renderer itself (see Training data). A scripted task-and-motion pipeline produces them in three steps.
Built on MolmoSpaces and CuRobo. Each episode places the target, the robot and synchronized multi-view cameras, then records RGBD, proprioception, actions and robot masks.
Each target carries a precomputed library of antipodal grasps. Candidates are ranked by reachability, alignment and approach stability, then collision-checked and passed through batched IK.
Task-specific primitives drive the motion: pregrasp–grasp–lift–place–retreat for pick-and-place, constrained handle motion for articulated objects. On IK loss or gripper slip the planner retries with a new grasp; only successful episodes are kept.
Every method starts from the same trajectories and turns them into training video; only the rendering differs. We compare against the following baselines:
Every pane is the same trajectory from the same camera, frame for frame, so only appearance changes.
One successful rollout per task, across all three embodiments. All clips play at 2× speed.
Mug
Droid
Bowl
Droid
Apple
Droid
Egg
Droid
Open the fridge
Droid
Open the drawer
Droid
Pick up the radio
R1Pro
Place the green marker
Yam
| Droid | R1Pro | Yam | |||||||
|---|---|---|---|---|---|---|---|---|---|
| Data | Mug | Bowl | Apple | Egg | Open Fridge | Open Drawer | Pick up Radio | Place Green Marker | Avg. |
| Sim | 40% | 30% | 10% | 0% | 0% | 0% | 0% | 0% | 10% |
| Sim + DR | 60% | 20% | 20% | 30% | 0% | 0% | 30% | 0% | 20% |
| RoboRender | 80% | 80% | 60% | 50% | 70% | 90% | 70% | 70% | 71% |
Each task is evaluated over 10 rollouts with varied robot starts and object locations. Pick-and-place tasks additionally include random table distractors.
| Training data | Drawer | Fridge |
|---|---|---|
| Wan-Fun-Control | 0% | 0% |
| Cosmos-Transfer2.5 | 30% | 0% |
| RoboRender | 90% | 70% |
Policies trained on the same simulation trajectories rendered by different video models. The ordering matches the generation-quality ranking above, so video quality translates directly into policy success.
| Training data | YAM | Droid PnP | R1Pro |
|---|---|---|---|
| Depth only | 30% | 25% | 10% |
| RoboRender | 70% | 68% | 70% |
Depth sidesteps the visual gap but discards semantics. Depth-only policies estimate depth from RGB at inference via VideoDepthAnything; RGB cues turn out to matter for identifying which object the task refers to.
The open extension on each Droid pick-and-place bar is the same method evaluated without distractors. Adding them degrades both baselines sharply — their training data has no clutter — while RoboRender barely moves, since it can insert distractors during generation. On articulated tasks it is the only method with non-zero success: these need fine visual detail, like drawer and fridge handles whose texture matches the object body.
Policy success rate against the number of videos generated per unique simulation trajectory, for pick-and-place (left) and opening (right). The solid dark red line is the per-category average, the dash-dot lines the individual tasks, and the dashed grey line a reference policy trained with 2× more simulation trajectories. Performance improves with more generated videos even though the underlying set of unique simulation trajectories is unchanged.