Agreement with ground-truth instance segmentation.
SceneRig
An Agentic System for Simulation-Ready 3D Scene Reconstruction from Single Images
1 Luma AI 2 Cornell University
SceneRig reconstructs individual objects and refines their placement using geometric and physical feedback.
01 / Reconstruction
Reconstruction comparisons
Source-view and nearby-view renders compare object geometry, appearance, and placement across SceneRig and reconstruction baselines. Click a reconstruction to explore it in 3D.
We test our system on diverse scenes from simulator images, real robot workspaces, smartphone photos, and internet or AI-generated images. We also simulate every reconstruction for five seconds to test whether its objects remain stable under gravity.
Loading reconstruction comparisons…
3D position error, including unmatched objects.
Scenes with every object passing the stability test.
SceneRig vs. VIGA, using Claude Opus 5 where applicable. Segmentation and placement: 44 simulator-grounded scenes. Stability: 80 scenes; every object must move less than 50 mm over five seconds of simulation. OSPA uses a 5 cm cutoff. VIGA includes object-grouping instructions.
02 / Robotics
Robotics applications
We evaluate whether reconstructed scenes predict real-robot policy outcomes and reproduce recorded manipulations through open-loop trajectory replay.
Real-to-sim policy evaluation
The same policy runs independently in the real workspace and its reconstruction. We compare outcomes at two levels.
We evaluate π₀.₅ and LingBot-VLA 2.0 on 100 real episodes across five manipulation tasks. For each episode, SceneRig and SimFoundry reconstruct the same starting workspace from a calibrated RGB-D observation, then run the policy checkpoint used on the real robot. Both simulations share the robot calibration, controllers, and physics settings, so the comparison tests how well each reconstructed scene predicts the real outcome.
| SceneRig | SimFoundry | |
|---|---|---|
| Episode-level agreementMatching real and simulated success/failure outcomes. | 69% | 54% |
| Task success-rate correlationPearson’s r across 10 policy–task pairs. | 0.92 | 0.84 |
Open-loop trajectory replay
Recorded joint and gripper commands are replayed in simulation without policy queries or trajectory adaptation. We measure task success after replay.
We use 25 successful teleoperated demonstrations, with five examples of each task. Both methods reconstruct the same initial workspace and replay identical commands with shared robot calibration, controllers, and physics settings. This tests whether the recovered geometry and object placement support the original manipulation, including grasping, moving, and placing objects into containers.
| Task | SceneRig | SimFoundry |
|---|---|---|
| All tasks | 80% |
36% |
| Everything → bin | 3/5 |
0/5 |
| Cup → bowl | 4/5 |
3/5 |
| Fruits → bowl | 4/5 |
2/5 |
| Marker → cup | 4/5 |
0/5 |
| Mustard → bin | 5/5 |
4/5 |
Recorded robot interactions
Representative videos illustrate individual episodes; the aggregate scores above include all evaluated episodes. For these robotics experiments, both methods reconstruct the same armless head-camera image using measured metric depth and calibrated intrinsics in place of monocular estimates. The calibrated head-camera pose registers each scene to the robot. Both methods share controllers and physics. Policy agreement and replay success measure different capabilities.
03 / Method
Method overview
SceneRig reconstructs individual objects, supporting surfaces, and their spatial relations from a single RGB image. Agents complete the scene geometry, appearance, and lighting in a persistent Blender scene. During object pose refinement, the agent identifies the object and pose component to adjust; the backend searches bounded updates and tests their geometric alignment and physical feasibility.
Relational and metric initialization
Recover objects, surfaces, and support relations; reconstruct individual meshes and initialize their metric placement. Settle objects in support order, refining alignment with observed geometry.
Agentic scene construction
Geometry, material, and lighting agents construct root surfaces and their appearance, with geometric checks and independent visual verification. Imported object geometry and appearance stay fixed.
Physics-grounded object pose refinement
Agents diagnose pose errors. A bounded backend searches adjustments and tests them in simulation. Complex fixes use direct scene-code edits, followed by settling and review.
View recorded tool calls03 / Pose refinement
Tool-call trace
Recorded investigation, pose updates, and backend feedback.
Loading the recorded pose-refinement trace…
USD export. A final joint gravity simulation sets object poses. The exported scene includes meshes, textures, lights, cameras, object identities, collision geometry, and physical properties.
.usd

