Blog

The Sim-to-Real Gap in Robotics: Why It Happens and How to Close It

Synthetic Data & Simulation
10 min

A grasping model hits 90% success in simulation. On the real arm, it misses half its picks. That drop is the sim-to-real gap, and almost every robotics team that trains in simulation runs into it sooner or later.

The sim-to-real gap is the difference in performance between a model or control policy evaluated in simulation and the same model running in the physical world. It exists because no simulator reproduces reality exactly: rendering, physics, sensors and actuators are all approximations. The gap can't be removed completely, but it can be measured and made small enough that simulation becomes a reliable place to train.

What Is the Sim-to-Real Gap?

Simulation is attractive for an obvious reason. You can run thousands of training episodes in parallel (NVIDIA Isaac Lab runs thousands of environments on a single GPU), you never break hardware, and every frame comes with free ground truth. The catch is that the model learns the simulator, not the world. Whatever the simulator gets wrong, the model treats as truth.

Researchers usually define the gap in concrete terms: take one evaluation metric, compute it in simulation and on the real system, and compare. A recent survey, The Reality Gap in Robotics, frames it this way and groups the main fixes into domain randomization, real-to-sim transfer, state and action abstractions, and sim-real co-training.

Perception models (detection, segmentation, pose estimation, depth) suffer mostly from visual differences. Images from a renderer don't look like images from a camera, and the model notices. Control policies (grasping, locomotion, manipulation) suffer mostly from dynamics: friction, contact, motor response, latency. A policy can see the world perfectly and still fail because the gripper closes 40 milliseconds later than it did in simulation. Most real systems have both problems at once, but knowing which one dominates tells you where to spend effort.

Where the Sim-to-Real Gap Comes From

NVIDIA's robotics documentation describes the gap as a combination of gaps in sensing, actuation, physics and modeling. In practice, it is useful to split the visual side further into appearance and content.

Appearance Gap: Rendering, Lighting and Materials

This is the pixel-level difference between a rendered image and a real photo of the same scene. Real surfaces have scratches, dust, fingerprints and subsurface scattering. Real lighting mixes daylight, fluorescent flicker and reflections off things outside the frame. Even good renderers tend to produce surfaces that are too clean and shadows that are too sharp.

A detector trained on clean renders can learn shortcuts, like a specific highlight on a metal part, that never appear on the factory floor.

Content Gap: What Your Simulated Scenes Leave Out

The content gap is about what is in the scene, not how it looks. Your simulated warehouse has 12 box types; the real one has 300, plus shrink wrap, torn labels and a coffee cup someone left on a shelf. Your simulated kitchen never has a towel half hanging off the counter.

The content gap is where long-tail failures live. A model can handle every object it saw in simulation and still fail on the first unusual configuration in production. This is why generating edge cases for computer vision models deliberately matters more than adding more of the same scenes.

Physics and Contact Dynamics

Physics engines approximate contact. Friction coefficients are constants; in reality they change with surface wear, humidity and load. Rigid objects are modeled reasonably well. Cloth, cables, food, liquids, granular material and anything soft are not. A simulated towel is a clean mesh. A real towel bunches, clings and slides out of the gripper.

For control policies, this is usually the biggest source of the gap.

Sensor and Actuator Mismatch

Simulated cameras are often ideal pinhole cameras: no noise, no motion blur, no lens distortion, no rolling shutter, perfect exposure. Simulated depth sensors return clean values on glass and black plastic, where real ones return holes. On the actuation side, motor models usually leave out backlash, friction, thermal drift and control latency.

Source of the gap Typical symptom Common mitigation
Appearance (lighting, materials, textures) Detector works on renders, drops sharply on camera images Domain randomization, physically based rendering, image translation
Content (objects, layouts, clutter) Failures on unseen object types or unusual arrangements Larger asset libraries, scene variation, targeted edge cases
Physics (contact, friction, deformables) Grasps slip, objects behave differently when pushed System identification, dynamics randomization
Sensors (noise, blur, depth artifacts) Unstable predictions, failures on reflective or dark surfaces Sensor noise models, augmentation, calibration against real captures
Actuators (latency, backlash) Overshoot, oscillation, timing errors Actuator modeling, latency randomization, real-world fine-tuning

How to Measure the Sim-to-Real Gap

If you don't put a number on the gap, you can't tell whether a change helped. The basic measurement is simple. Pick one metric that matters for the task (mAP for detection, IoU for segmentation, success rate for grasping), evaluate the same model on a held-out simulated set and on a real test set, and track the difference. The real test set doesn't need to be large. A few hundred carefully labeled real images, or 50 to 100 real grasp attempts, already tell you a lot. What matters is that it stays fixed, so every experiment is compared against the same reference.

For perception models, go one step further and split the error. Run the model on real images and look at which failures dominate. If it misses objects it knows under unusual lighting, you have an appearance problem. If it fails on objects or arrangements it never saw, you have a content problem. This diagnosis decides whether to invest in rendering or in scene variety, and teams that skip it often spend months improving the wrong one.

The same evaluation set also tells you when synthetic data itself is the problem. We cover the checks in more detail in our guide on how to validate synthetic data quality before model training.

How to Close the Sim-to-Real Gap

No single technique closes the gap. The right mix depends on whether it is mostly visual or mostly physical.

Domain Randomization and Structured Randomization

Domain randomization flips the problem around. Instead of making simulation match reality, you make simulation so varied that reality looks like just one more variation. Textures, lighting, camera positions, object colors and distractor objects all change randomly from frame to frame, and the model learns to ignore them.

The idea became well known after Tobin et al. (2017) trained an object localization model only on randomized renders and used it on a real robot. OpenAI later applied the same principle to physics parameters for its robotic hand that solved a Rubik's cube.

Pure randomization has a cost, though. Fully random scenes waste training data on configurations that never occur. Structured randomization keeps the variation inside realistic bounds: boxes sit on shelves, not floating in the air, and lighting follows plausible setups for the environment. NVIDIA's example with Trimble used structured randomization in Isaac Replicator to train a detector that found doors in real buildings using synthetic indoor scenes only.

Higher-Fidelity Rendering and Sensor Models

The opposite approach narrows the appearance gap directly. Physically based materials, ray-traced lighting, realistic camera models with noise, blur and distortion, and depth sensors that fail where real ones fail. This costs more per frame and more effort to build, but it reduces how much randomization the model needs.

In practice, the two work together: realistic assets as the baseline, randomization on top for what you can't predict. We compare rendering-based and learned approaches in more depth in world models vs 3D simulation for Physical AI.

System Identification for Robot Dynamics

For control policies, the first step is usually to make the simulated robot match the real one. System identification means measuring the real robot's parameters (joint friction, motor response, mass distribution, latency) and fitting the simulator to them. Teams then randomize those parameters around the measured values, so the policy tolerates the remaining error.

Real-to-Sim: Digital Twins and Scene Reconstruction

Real-to-sim works in the opposite direction from sim-to-real. You capture the real environment, with photogrammetry, 3D scanning or neural reconstruction methods like Gaussian splatting, and rebuild it in simulation. The model then trains in a virtual copy of the place where it will actually work.

This works best for fixed environments such as a specific production line or warehouse. It's less useful when the robot will face many different sites, since one reconstructed scene can make the model overfit to it.

Domain Adaptation and Image Translation

Domain adaptation changes the data or the model so simulated and real inputs look alike to the network. At the pixel level, generative models translate rendered images into a realistic style. At the feature level, training pushes the network to produce the same internal representation for sim and real images. Google's GraspGAN work combined both to train grasping with far fewer real examples.

These methods need at least some unlabeled real data, and image translation can quietly move object boundaries, which breaks labels. Check the translated images before trusting them.

Sim-Real Co-Training and Fine-Tuning on Real Data

The most reliable method in practice is also the least exotic: train mostly on synthetic data and mix in a small amount of real data. Either fine-tune a model pretrained in simulation on real samples, or train on both sets at once with a fixed ratio. Synthetic data provides scale, labels and rare cases. Real data anchors the model to what the sensors actually produce. Our article on synthetic data for robotics training covers how teams typically balance the two.

Why Domain Randomization Alone Is Not Enough

Domain randomization is often presented as the answer to the sim-to-real gap. It solves part of it. Randomization covers variation you thought of. If you randomize lighting and textures but every object in your asset library is rigid, no amount of randomization teaches the model what a crumpled bag looks like. That's a content gap, and randomizing it away requires the assets to exist first.

It also struggles with physics. You can randomize friction between 0.3 and 1.2, but if the simulator can't model deformable contact at all, the policy never sees the behavior it will meet in the real world. And heavy randomization has a price in model capacity: a network that must be robust to everything may end up less accurate on the conditions that actually occur.

A Practical Sim-to-Real Workflow for Perception Models

For teams training detection or segmentation models on synthetic data, this sequence covers the main steps:

  1. Collect a small, fixed real test set from the deployment environment and label it carefully. Every later decision is measured against it.
  2. Build or source 3D assets that match the real objects in geometry and materials, including variants such as worn, dirty or damaged versions.
  3. Generate a baseline synthetic dataset with structured randomization of lighting, camera pose, backgrounds and object placement.
  4. Train, evaluate on the real test set, and sort failures into appearance and content problems.
  5. Fix the dominant problem first: better materials and sensor models for appearance, more objects, layouts and edge cases for content.
  6. Add a small amount of real labeled data through fine-tuning or co-training, then measure the gap again.
  7. Repeat whenever the deployment environment changes, such as a new site, new products or a new camera.

The order matters. Without a real test set from day one, you can't tell which of your synthetic images helped.

Where Synthetic Data Fits

Synthetic data doesn't eliminate the sim-to-real gap. It makes the gap something you can control. Because every image comes from a known scene, you can change one factor at a time, generate the rare cases real data collection never captures, and get labels without manual annotation.

Vivid 3D builds synthetic datasets for robotics and computer vision from controlled 3D scenes, with variation in lighting, materials, camera placement and object configuration, and annotations generated from the scene itself. If your model performs well in simulation but not on the robot, see how we work with robotics and Physical AI teams.

Closing the sim-to-real gap is less about one breakthrough technique and more about discipline: measure the gap on real data, find out which part of it dominates, and fix that part first.

Frequently Asked Questions

What causes the sim-to-real gap?

The gap comes from differences between the simulator and the real world in four areas: visual appearance (rendering, lighting, materials), scene content (objects and layouts the simulation doesn't include), physics (contact, friction, deformable objects) and hardware (sensor noise and actuator behavior). Perception models are mostly affected by the first two, control policies by the last two.

Can domain randomization eliminate the sim-to-real gap?

No. Domain randomization makes models robust to variation you include in the simulation, but it can't cover objects, materials or physical behavior the simulator doesn't model. Most teams combine it with realistic assets, system identification and a small amount of real data.

How much real data do you need to close the gap?

There is no fixed number, since it depends on the task and how close your simulation already is. For evaluation, a few hundred labeled real images or 50 to 100 real trials are often enough to see the gap clearly. For training, the real set is typically a small fraction of the synthetic one.

What is the difference between sim-to-real transfer and domain adaptation?

Sim-to-real transfer is the overall goal: making a model trained in simulation work in the real world. Domain adaptation is one family of methods for reaching that goal, which align simulated and real data at the pixel or feature level so the model treats them the same way.

Is the sim-to-real gap smaller for perception than for control?

Usually, yes. Visual differences can be reduced with better rendering, randomization and some real images, so perception models can often train mostly on synthetic data. Control policies also face contact physics and actuator dynamics, which current simulators model less accurately, so their gap tends to be harder to close.

Table of contents
10 min
Share

Recommended