RoboRender: Robot-Oriented Video Generation for Visual Sim-to-Real Transfer

Huang Huang1*, Wensi Ai1*, Ziyu Chen1*, Youhui Wang1*, Zijian Du2, Yang Liu3, Jiaolong Yang3, Li Fei-Fei1, Jiajun Wu1
1Stanford University 2Nvidia 3Microsoft Research
*Equal contribution

From simulation trajectories, RoboRender generates diverse and photorealistic videos oriented for visual sim-to-real robot policy transfer, supporting multi-view consistency, distractor insertion, and cross-embodiment generation.

Abstract

Simulation enables large-scale, low-cost robot data generation, but policies trained in simulation often fail to transfer to the real world due to the sim-to-real visual discrepancies. Existing approaches often rely on intermediate representations, which can discard rich semantic information or require additional perception modules at deployment. We address this visual sim-to-real gap with RoboRender, a framework that converts simulated trajectories into photorealistic RGB videos for policy learning. RoboRender trains a robot-oriented video generation model conditioned on simulated depth videos, language instructions, and robot RGB mask videos, preserving simulator geometry, robot motion, and action labels while synthesizing realistic textures, backgrounds, and distractors. The generated RGB videos are paired with simulator-provided states and actions to train policies for zero-shot real-world deployment. On robot video test sets, our video model outperforms depth-conditioned video generation baselines in generation quality. In real-world experiments across pick-and-place, articulated-object manipulation, and mobile manipulation tasks, policies trained on RoboRender-generated data achieve a 71% average success rate, outperforming raw simulation renderings and conventional visual domain randomization by approximately 7.1× and 3.6×, respectively. We further show that policy performance improves with more generated videos per simulation trajectory, increasing opening-task success by 65 percentage points. These results demonstrate that generative video rendering mitigates the visual sim-to-real gap for zero-shot policy transfer.

Video Model

RoboRender adapts Wan2.1-T2V-1.3B into a robot-oriented renderer, fine-tuned on ~130k real robot manipulation samples (~70k from AgiBot-World-Alpha, ~60k from DROID) at 416×240 — see Training data for the corpus and how it was built. The model is trained on real video and only conditioned on simulation at generation time.

RoboRender video generation model

What the model can do

  • Multi-view consistency. Views are stacked into a unified multi-view latent rather than generated independently, so shared objects stay consistent across the external and wrist cameras.
  • Long-horizon rollouts. Generation is autoregressive: each chunk is conditioned on the previous chunk's last frame, so variable-length demonstrations can exceed the model's fixed frame window.
  • Cross-embodiment generation. Robot-mask conditioning lets the model render embodiments it never saw in training — both the R1Pro used for mobile manipulation and the YAM arm are unseen during video-model training.

Rendering cost

Rendering a whole simulation dataset takes many diffusion passes, so inference cost bounds how much data you can generate. Two optimizations bring it down:

  • TeaCache — reuses intermediate diffusion features across denoising steps. A 3-view 81-frame clip drops from 68.8 s to 35.3 s on a single H200, a 1.95× speedup at negligible quality loss.
  • DMD2 step distillation — reduces the 50-step bidirectional sampler to 6 steps. The distilled model renders a 3-view clip in 4.29 s.

Sample generated episodes

Real-Robot Experiments

Every policy here is trained purely in simulation and deployed zero-shot — no real-world fine-tuning, and no real data beyond what trained the renderer. Each is a fine-tuned π0.5 predicting 32-step action chunks with the first 16 executed, trained for 40k steps at batch size 72. Evaluation runs in a natural office environment, where object size, geometry, camera viewpoint and background all differ from simulation, across eight tasks and three embodiments.

Trajectory generation

The policies are trained on simulation trajectories, which are separate from the real corpus that trained the renderer itself (see Training data). A scripted task-and-motion pipeline produces them in three steps.

Scene setup

Built on MolmoSpaces and CuRobo. Each episode places the target, the robot and synchronized multi-view cameras, then records RGBD, proprioception, actions and robot masks.

Grasp selection

Each target carries a precomputed library of antipodal grasps. Candidates are ranked by reachability, alignment and approach stability, then collision-checked and passed through batched IK.

Planning and retries

Task-specific primitives drive the motion: pregrasp–grasp–lift–place–retreat for pick-and-place, constrained handle motion for articulated objects. On IK loss or gripper slip the planner retries with a new grasp; only successful episodes are kept.

Baseline comparison

Every method starts from the same trajectories and turns them into training video; only the rendering differs. We compare against the following baselines:

  • Simulation — the simulator's own render. Clean geometry, but synthetic textures, flat lighting and an unrealistic background.
  • Simulation + DR — domain randomization over textures and colors. Widens the distribution, but the result is visually incoherent and still does not match real material and lighting statistics.
  • Depth — the simulated depth map. This is the conditioning signal our video model consumes, but we additionally train a policy on depth alone, to show how much the RGB appearance matters (see Against depth-only policies).
  • RoboRender (ours) — frames generated from the same trajectory. Realistic surfaces, reflections and lighting, with the simulator's geometry preserved.

Every pane is the same trajectory from the same camera, frame for frame, so only appearance changes.

Sample RoboRender policy rollouts

One successful rollout per task, across all three embodiments. All clips play at 2× speed.

Mug
Droid

Bowl
Droid

Apple
Droid

Egg
Droid

Open the fridge
Droid

Open the drawer
Droid

Pick up the radio
R1Pro

Place the green marker
Yam

Real-world success rate across tasks

Droid R1Pro Yam
Data Mug Bowl Apple Egg Open Fridge Open Drawer Pick up Radio Place Green Marker Avg.
Sim 40%30%10%0%0%0%0%0% 10%
Sim + DR 60%20%20%30%0%0%30%0% 20%
RoboRender 80%80%60%50% 70%90% 70%70% 71%

Each task is evaluated over 10 rollouts with varied robot starts and object locations. Pick-and-place tasks additionally include random table distractors.

Against video-model baselines

Training dataDrawerFridge
Wan-Fun-Control0%0%
Cosmos-Transfer2.530%0%
RoboRender90%70%

Policies trained on the same simulation trajectories rendered by different video models. The ordering matches the generation-quality ranking above, so video quality translates directly into policy success.

Against depth-only policies

Training dataYAMDroid PnPR1Pro
Depth only30%25%10%
RoboRender70%68%70%

Depth sidesteps the visual gap but discards semantics. Depth-only policies estimate depth from RGB at inference via VideoDepthAnything; RGB cues turn out to matter for identifying which object the task refers to.

Success rate by task category

    The open extension on each Droid pick-and-place bar is the same method evaluated without distractors. Adding them degrades both baselines sharply — their training data has no clutter — while RoboRender barely moves, since it can insert distractors during generation. On articulated tasks it is the only method with non-zero success: these need fine visual detail, like drawer and fridge handles whose texture matches the object body.

    Scaling generated videos per trajectory

    Policy success rate against the number of videos generated per unique simulation trajectory, for pick-and-place (left) and opening (right). The solid dark red line is the per-category average, the dash-dot lines the individual tasks, and the dashed grey line a reference policy trained with 2× more simulation trajectories. Performance improves with more generated videos even though the underlying set of unique simulation trajectories is unchanged.