SceneRig

An Agentic System for Simulation-Ready 3D Scene Reconstruction from Single Images

Rundong Luo1,2 Yunong Liu1 Di Cao1 Matthew Tancik1 Shyamal Buch1 Colton Stearns1 Wenqi Xian1

1 Luma AI 2 Cornell University

SceneRig reconstructs individual objects and refines their placement using geometric and physical feedback.

01Input imageSingle-view RGB observation
Input photograph of a coffee workspace with mugs, a coffee machine, and a kettle.Input
02Reconstruction closeupsObject geometry & placement
SimFoundry reconstruction closeup: the rear black mug lies on its side.SimFoundry
SceneRig reconstruction closeup: the rear black mug remains upright, matching the input.SceneRig
033D scene under gravityFive seconds of simulation · RGB & clay
Gravity on
Loading 0.0 / 5.0 s

Recorded physics · 1× speed · drag the divider to compare materials.

Recording details ↗
04Robotics applicationsExperiments in reconstructed robot workspaces
Robot experiments use separate input images

Loading recorded robot demonstrations…

01 / Reconstruction

Reconstruction comparisons

Source-view and nearby-view renders compare object geometry, appearance, and placement across SceneRig and reconstruction baselines. Click a reconstruction to explore it in 3D.

We test our system on diverse scenes from simulator images, real robot workspaces, smartphone photos, and internet or AI-generated images. We also simulate every reconstruction for five seconds to test whether its objects remain stable under gravity.

Reconstruction metrics

Instance segmentation ARI ↑
0.878 vs. 0.283

Agreement with ground-truth instance segmentation.

3D object placement OSPA@5 ↓
2.4 cm vs. 4.9 cm

3D position error, including unmatched objects.

Scene stability Pass rate ↑
100% vs. 10.0%

Scenes with every object passing the stability test.

SceneRig vs. VIGA, using Claude Opus 5 where applicable. Segmentation and placement: 44 simulator-grounded scenes. Stability: 80 scenes; every object must move less than 50 mm over five seconds of simulation. OSPA uses a 5 cm cutoff. VIGA includes object-grouping instructions.

02 / Robotics

Robotics applications

We evaluate whether reconstructed scenes predict real-robot policy outcomes and reproduce recorded manipulations through open-loop trajectory replay.

A

Real-to-sim policy evaluation

The same policy runs independently in the real workspace and its reconstruction. We compare outcomes at two levels.

We evaluate π₀.₅ and LingBot-VLA 2.0 on 100 real episodes across five manipulation tasks. For each episode, SceneRig and SimFoundry reconstruct the same starting workspace from a calibrated RGB-D observation, then run the policy checkpoint used on the real robot. Both simulations share the robot calibration, controllers, and physics settings, so the comparison tests how well each reconstructed scene predicts the real outcome.

SceneRigSimFoundry
Episode-level agreementMatching real and simulated success/failure outcomes. 69%54%
Task success-rate correlationPearson’s r across 10 policy–task pairs. 0.920.84
B

Open-loop trajectory replay

Recorded joint and gripper commands are replayed in simulation without policy queries or trajectory adaptation. We measure task success after replay.

We use 25 successful teleoperated demonstrations, with five examples of each task. Both methods reconstruct the same initial workspace and replay identical commands with shared robot calibration, controllers, and physics settings. This tests whether the recovered geometry and object placement support the original manipulation, including grasping, moving, and placing objects into containers.

TaskSceneRigSimFoundry
All tasks
80%
36%
Everything → bin
3/5
0/5
Cup → bowl
4/5
3/5
Fruits → bowl
4/5
2/5
Marker → cup
4/5
0/5
Mustard → bin
5/5
4/5

Task-level comparison

Real and simulated success rates

Real versus simulated task success rates for π₀.₅ and LingBot-VLA 2.0. SceneRig has pooled Pearson correlation r = 0.92; SimFoundry has r = 0.84. Shapes distinguish policies and colors distinguish tasks.
Each point represents one policy on one task, with 10 episodes per task. Pearson’s r pools all 10 policy–task pairs; it measures how success rates vary together, rather than how often individual episode outcomes match. The dashed line indicates equal real and simulated success rates. Nested markers show overlapping results.

Recorded robot interactions

Representative videos illustrate individual episodes; the aggregate scores above include all evaluated episodes. For these robotics experiments, both methods reconstruct the same armless head-camera image using measured metric depth and calibrated intrinsics in place of monocular estimates. The calibrated head-camera pose registers each scene to the robot. Both methods share controllers and physics. Policy agreement and replay success measure different capabilities.

03 / Method

Method overview

SceneRig reconstructs individual objects, supporting surfaces, and their spatial relations from a single RGB image. Agents complete the scene geometry, appearance, and lighting in a persistent Blender scene. During object pose refinement, the agent identifies the object and pose component to adjust; the backend searches bounded updates and tests their geometric alignment and physical feasibility.

Reconstructed objects physically settled on their recovered supports under gravity.
01 / Initialization

Relational and metric initialization

Recover objects, surfaces, and support relations; reconstruct individual meshes and initialize their metric placement. Settle objects in support order, refining alignment with observed geometry.

The persistent Blender scene after root-surface appearance and lighting refinement.
02 / Scene construction

Agentic scene construction

Geometry, material, and lighting agents construct root surfaces and their appearance, with geometric checks and independent visual verification. Imported object geometry and appearance stay fixed.

Paired object crops inspected by the pose agent to diagnose a spoon placement error.
03 / Pose refinement

Physics-grounded object pose refinement

Agents diagnose pose errors. A bounded backend searches adjustments and tests them in simulation. Complex fixes use direct scene-code edits, followed by settling and review.

View recorded tool calls

03 / Pose refinement

Tool-call trace

Recorded investigation, pose updates, and backend feedback.

Loading the recorded pose-refinement trace…

USD export. A final joint gravity simulation sets object poses. The exported scene includes meshes, textures, lights, cameras, object identities, collision geometry, and physical properties.

.usd