Real-to-Simulation Scene Foundry
A modular video-to-simulation pipeline for reconstructing physics-ready robot environments.
Original project: SimFoundry · NVIDIA Research and research collaborators

01 / Overview
The system at a glance.
SimFoundry reconstructs simulation-ready scenes from short real-world videos, then supports scene variation and application in OmniGibson. Its modular pipeline combines video processing, segmentation, depth, object decomposition, 3D generation, pose estimation, physical parameter compilation, and USD export.
- Original project
- SimFoundry
- Created by
- NVIDIA Research and research collaborators
- License
- Apache-2.0 for NVIDIA-owned source; integrated models, datasets, SDKs, and derived components may use separate terms
- Source review
- August 22, 2026
02 / Challenge
The engineering problem.
Robot learning needs diverse, physically meaningful environments, but manually rebuilding real scenes in simulation is slow and specialized.
SimFoundry explores how foundation models and explicit physics compilation can automate much of that conversion while retaining modular stages that can evolve independently.
03 / System design
How the architecture responds.
Pipeline A performs a 13-stage reconstruction and exports a USD/OmniGibson scene. Pipeline B produces digital-cousin variations and task proposals. Pipeline C loads results into OmniGibson for smoke testing, teleoperation data collection, and policy evaluation. Hydra configuration, staged environments, GPU budgeting, and validation scripts make the research stack operable as a system rather than a single model demo.
- 01
Capture a short scene video
- 02
Estimate depth and segment the environment
- 03
Decompose objects and generate 3D meshes
- 04
Recover poses and physical parameters
- 05
Compile and validate a USD scene
- 06
Generate variants or evaluate in OmniGibson
04 / Technology
The implementation stack.
A concise, repository-backed view of the primary platforms, protocols, models, and runtime tools.
- Python
- CUDA
- OmniGibson
- USD
- Hunyuan3D-2
- Depth Anything 3
- SAM3
- FoundationPose
- Hydra
- Gemini / Vertex AI
- FFmpeg
05 / Highlights
What makes the system notable.
- Thirteen-stage reconstruction pipeline
- Modular reconstruction, augmentation, and application layers
- GPU-budget-aware execution
- Physics-ready scene compilation
- Digital-cousin scene variation
06 / Repository evidence
Verifiable signals.
These statements are derived from the project’s current README, source structure, or license—not from NexLoomix client work.
- The current release documents Linux, NVIDIA CUDA, FFmpeg, substantial local storage, Hugging Face access, and Gemini or Vertex AI access as requirements.
- Reconstruction covers video processing through USD/OmniGibson export across 13 stages.
- The repository marks example assets and some broader generation, training, and evaluation material as still forthcoming in the initial release.
07 / NexLoomix perspective
The transferable product lesson.
The architecture is a useful blueprint for complex AI pipelines: separate reconstruction, augmentation, and application so each stage can be tested, replaced, and governed independently.
This interpretation is editorial analysis by NexLoomix. It is intentionally separated from the verified repository evidence above.
08 / Source & attribution
Credit where it belongs.
This is presented as a research-engineering reference. Several optional dependencies are non-commercial, research-only, gated, or otherwise restricted.
- Repository
- SimFoundry on GitHub
- Author / organization
- NVIDIA Research and research collaborators
- License summary
- Apache-2.0 for NVIDIA-owned source; integrated models, datasets, SDKs, and derived components may use separate terms · review license
- NexLoomix relationship
- Independent editorial showcase; no authorship, client relationship, partnership, or endorsement claimed.
09 / Related NexLoomix services
Build a system around the lesson.
These capabilities connect the reference architecture to a new product shaped around your own users, data, constraints, and ownership.
From reference to product
Build an intelligent system around your business.
Bring us the workflow, constraint, or opportunity. We’ll help shape the right product and architecture.