

A grasping model hits 90% success in simulation. On the real arm, it misses half its picks. That drop is the sim-to-real gap, and almost every robotics team that trains in simulation runs into it sooner or later.
The sim-to-real gap is the difference in performance between a model or control policy evaluated in simulation and the same model running in the physical world. It exists because no simulator reproduces reality exactly: rendering, physics, sensors and actuators are all approximations. The gap can't be removed completely, but it can be measured and made small enough that simulation becomes a reliable place to train.
Simulation is attractive for an obvious reason. You can run thousands of training episodes in parallel (NVIDIA Isaac Lab runs thousands of environments on a single GPU), you never break hardware, and every frame comes with free ground truth. The catch is that the model learns the simulator, not the world. Whatever the simulator gets wrong, the model treats as truth.
Researchers usually define the gap in concrete terms: take one evaluation metric, compute it in simulation and on the real system, and compare. A recent survey, The Reality Gap in Robotics, frames it this way and groups the main fixes into domain randomization, real-to-sim transfer, state and action abstractions, and sim-real co-training.
Perception models (detection, segmentation, pose estimation, depth) suffer mostly from visual differences. Images from a renderer don't look like images from a camera, and the model notices. Control policies (grasping, locomotion, manipulation) suffer mostly from dynamics: friction, contact, motor response, latency. A policy can see the world perfectly and still fail because the gripper closes 40 milliseconds later than it did in simulation. Most real systems have both problems at once, but knowing which one dominates tells you where to spend effort.
NVIDIA's robotics documentation describes the gap as a combination of gaps in sensing, actuation, physics and modeling. In practice, it is useful to split the visual side further into appearance and content.
This is the pixel-level difference between a rendered image and a real photo of the same scene. Real surfaces have scratches, dust, fingerprints and subsurface scattering. Real lighting mixes daylight, fluorescent flicker and reflections off things outside the frame. Even good renderers tend to produce surfaces that are too clean and shadows that are too sharp.
A detector trained on clean renders can learn shortcuts, like a specific highlight on a metal part, that never appear on the factory floor.
The content gap is about what is in the scene, not how it looks. Your simulated warehouse has 12 box types; the real one has 300, plus shrink wrap, torn labels and a coffee cup someone left on a shelf. Your simulated kitchen never has a towel half hanging off the counter.
The content gap is where long-tail failures live. A model can handle every object it saw in simulation and still fail on the first unusual configuration in production. This is why generating edge cases for computer vision models deliberately matters more than adding more of the same scenes.
Physics engines approximate contact. Friction coefficients are constants; in reality they change with surface wear, humidity and load. Rigid objects are modeled reasonably well. Cloth, cables, food, liquids, granular material and anything soft are not. A simulated towel is a clean mesh. A real towel bunches, clings and slides out of the gripper.
For control policies, this is usually the biggest source of the gap.
Simulated cameras are often ideal pinhole cameras: no noise, no motion blur, no lens distortion, no rolling shutter, perfect exposure. Simulated depth sensors return clean values on glass and black plastic, where real ones return holes. On the actuation side, motor models usually leave out backlash, friction, thermal drift and control latency.
If you don't put a number on the gap, you can't tell whether a change helped. The basic measurement is simple. Pick one metric that matters for the task (mAP for detection, IoU for segmentation, success rate for grasping), evaluate the same model on a held-out simulated set and on a real test set, and track the difference. The real test set doesn't need to be large. A few hundred carefully labeled real images, or 50 to 100 real grasp attempts, already tell you a lot. What matters is that it stays fixed, so every experiment is compared against the same reference.
For perception models, go one step further and split the error. Run the model on real images and look at which failures dominate. If it misses objects it knows under unusual lighting, you have an appearance problem. If it fails on objects or arrangements it never saw, you have a content problem. This diagnosis decides whether to invest in rendering or in scene variety, and teams that skip it often spend months improving the wrong one.
The same evaluation set also tells you when synthetic data itself is the problem. We cover the checks in more detail in our guide on how to validate synthetic data quality before model training.
No single technique closes the gap. The right mix depends on whether it is mostly visual or mostly physical.
Domain randomization flips the problem around. Instead of making simulation match reality, you make simulation so varied that reality looks like just one more variation. Textures, lighting, camera positions, object colors and distractor objects all change randomly from frame to frame, and the model learns to ignore them.
The idea became well known after Tobin et al. (2017) trained an object localization model only on randomized renders and used it on a real robot. OpenAI later applied the same principle to physics parameters for its robotic hand that solved a Rubik's cube.
Pure randomization has a cost, though. Fully random scenes waste training data on configurations that never occur. Structured randomization keeps the variation inside realistic bounds: boxes sit on shelves, not floating in the air, and lighting follows plausible setups for the environment. NVIDIA's example with Trimble used structured randomization in Isaac Replicator to train a detector that found doors in real buildings using synthetic indoor scenes only.
The opposite approach narrows the appearance gap directly. Physically based materials, ray-traced lighting, realistic camera models with noise, blur and distortion, and depth sensors that fail where real ones fail. This costs more per frame and more effort to build, but it reduces how much randomization the model needs.
In practice, the two work together: realistic assets as the baseline, randomization on top for what you can't predict. We compare rendering-based and learned approaches in more depth in world models vs 3D simulation for Physical AI.
For control policies, the first step is usually to make the simulated robot match the real one. System identification means measuring the real robot's parameters (joint friction, motor response, mass distribution, latency) and fitting the simulator to them. Teams then randomize those parameters around the measured values, so the policy tolerates the remaining error.
Real-to-sim works in the opposite direction from sim-to-real. You capture the real environment, with photogrammetry, 3D scanning or neural reconstruction methods like Gaussian splatting, and rebuild it in simulation. The model then trains in a virtual copy of the place where it will actually work.
This works best for fixed environments such as a specific production line or warehouse. It's less useful when the robot will face many different sites, since one reconstructed scene can make the model overfit to it.
Domain adaptation changes the data or the model so simulated and real inputs look alike to the network. At the pixel level, generative models translate rendered images into a realistic style. At the feature level, training pushes the network to produce the same internal representation for sim and real images. Google's GraspGAN work combined both to train grasping with far fewer real examples.
These methods need at least some unlabeled real data, and image translation can quietly move object boundaries, which breaks labels. Check the translated images before trusting them.
The most reliable method in practice is also the least exotic: train mostly on synthetic data and mix in a small amount of real data. Either fine-tune a model pretrained in simulation on real samples, or train on both sets at once with a fixed ratio. Synthetic data provides scale, labels and rare cases. Real data anchors the model to what the sensors actually produce. Our article on synthetic data for robotics training covers how teams typically balance the two.
Domain randomization is often presented as the answer to the sim-to-real gap. It solves part of it. Randomization covers variation you thought of. If you randomize lighting and textures but every object in your asset library is rigid, no amount of randomization teaches the model what a crumpled bag looks like. That's a content gap, and randomizing it away requires the assets to exist first.
It also struggles with physics. You can randomize friction between 0.3 and 1.2, but if the simulator can't model deformable contact at all, the policy never sees the behavior it will meet in the real world. And heavy randomization has a price in model capacity: a network that must be robust to everything may end up less accurate on the conditions that actually occur.
For teams training detection or segmentation models on synthetic data, this sequence covers the main steps:
The order matters. Without a real test set from day one, you can't tell which of your synthetic images helped.
Synthetic data doesn't eliminate the sim-to-real gap. It makes the gap something you can control. Because every image comes from a known scene, you can change one factor at a time, generate the rare cases real data collection never captures, and get labels without manual annotation.
Vivid 3D builds synthetic datasets for robotics and computer vision from controlled 3D scenes, with variation in lighting, materials, camera placement and object configuration, and annotations generated from the scene itself. If your model performs well in simulation but not on the robot, see how we work with robotics and Physical AI teams.
Closing the sim-to-real gap is less about one breakthrough technique and more about discipline: measure the gap on real data, find out which part of it dominates, and fix that part first.
The gap comes from differences between the simulator and the real world in four areas: visual appearance (rendering, lighting, materials), scene content (objects and layouts the simulation doesn't include), physics (contact, friction, deformable objects) and hardware (sensor noise and actuator behavior). Perception models are mostly affected by the first two, control policies by the last two.
No. Domain randomization makes models robust to variation you include in the simulation, but it can't cover objects, materials or physical behavior the simulator doesn't model. Most teams combine it with realistic assets, system identification and a small amount of real data.
There is no fixed number, since it depends on the task and how close your simulation already is. For evaluation, a few hundred labeled real images or 50 to 100 real trials are often enough to see the gap clearly. For training, the real set is typically a small fraction of the synthetic one.
Sim-to-real transfer is the overall goal: making a model trained in simulation work in the real world. Domain adaptation is one family of methods for reaching that goal, which align simulated and real data at the pixel or feature level so the model treats them the same way.
Usually, yes. Visual differences can be reduced with better rendering, randomization and some real images, so perception models can often train mostly on synthetic data. Control policies also face contact physics and actuator dynamics, which current simulators model less accurately, so their gap tends to be harder to close.
