SceneMosaic: Fast, Diverse 3D Scene Generation from a Single Image
SceneMosaic generates diverse, physically valid 3D room layouts from a single image by combining image priors for speed, agent-based layout evolution for accuracy, and local patch decomposition for combinatorial diversity, achieving 24x speedup over baselines with zero collision rate.
SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation
Robots entering everyday homes require massive simulation environments for training and evaluation, but existing simulators suffer from sparse furniture and monotonous layouts. Real rooms are cluttered, with objects supporting and occluding each other. Two mainstream approaches exist: agent-based text-to-3D generation (e.g., SAGE) and parametric image-to-3D pipelines. Both hit walls: agent methods are slow (SAGE takes 7.56 hours per scene), while parametric methods are fast but physically implausible (20–26% collision rate). Moreover, neither produces truly diverse, coherent layout variations that reflect real-world rearrangements.
Method Overview
SceneMosaic, from Bo Dai's team at HKU, introduces a three-stage pipeline: (1) scene reconstruction and structuring, (2) agent-based layout evolution, and (3) diversity-driven scene composition. The core metaphor is a mosaic: a scene is not a fixed whole but an assembly of independent, replaceable local patches.
Stage 1: Scene Reconstruction and Structuring
A perception agent performs instance-level object registration. SAM3 segments and SAM3D independently reconstructs each object's mesh and initial 6D pose. A directed acyclic graph (DAG) schedules specialized sub-agents to infer room boundaries and object dependencies — attach (support) and contain (containment) — organizing the scene into a hierarchical scene tree.
Crucially, a local unit is defined as a non-leaf anchor node together with its directly supported children (e.g., a bookshelf and its books, a desk with monitor and keyboard). Local units are largely independent and can evolve separately. Sparse cross-unit functional constraints (e.g., chair facing desk) are explicitly extracted to avoid over-constraining diversity.
Stage 2: Agent-Based Layout Evolution
Before agent intervention, physics-based correction runs: containment adjustment and gravity simulation let objects resolve collisions and eliminate floating. Gravity direction is anchor-dependent: downward for floor anchors, upward for ceiling anchors, along wall normals for wall anchors, so adhered objects naturally settle onto supporting surfaces.
A Critic-Actor loop then refines layouts:
Critic reads orthographic renderings, collision reports, and history memory, outputting qualitative suggestions without coordinates (e.g., "move chair toward desk") to avoid unreliable numeric estimates.
Actor translates suggestions into symbolic pose expressions referencing the current layout state, evaluated by a sandbox evaluator for precise alignment, symmetry, and spacing.
Three safeguards prevent deadlocks: strict context scoping, reducing 3D pose adjustments to orthographic 2D translation and rotation, and an explicit layout memory recording repeated failure patterns.
Stage 3: Diversity-Driven Scene Composition
Because local layouts are expressed in their anchor's coordinate system, any child variant can combine with any parent variant. Taking the Cartesian product along the scene tree yields a massive candidate pool; one evolution run produces many variants at near-zero marginal cost. A novelty distance metric (combining relative position, absolute distance, and rotation difference) paired with dynamic Max-Min greedy search selects a compact, diverse set of representative scenes that reflect real-world layout changes.
Experiments on SceneEval-100
Evaluation uses semantic plausibility (POS, ROT) and physical validity (navigability NAV, collision rate COL, out-of-bounds rate OOB).
Quantitative Results
SceneMosaic matches the strongest baseline in semantic quality (POS 84.3 vs. 83.5) while drastically reducing physical violations. It achieves a 24× speedup over SceneSmith, generating a variant in only 0.03 hours — faster than all baselines.
Qualitative Comparison
SAM3D leverages visual priors for fast initialization but often yields physically implausible poses. SceneMosaic preserves the image's global spatial structure and then improves object relations via layout evolution, ensuring chairs rest on floors and books are properly supported by shelves. Baselines still exhibit frequent floating and interpenetration.
Ablation Studies
Removing the image prior increases generation time from 0.14 to 1.25 hours. Removing physics processing spikes collision rate from 0.0 to 18.9%. Switching to perspective 3D representation doubles inference time and degrades quality. Most critically, replacing local decomposition with global evolution loses the efficiency advantage: variant generation becomes 21× slower than SceneMosaic.
Diversity Evaluation
Replacing greedy search with random sampling yields more similar layouts, while repeated sampling is least diverse and slowest. Compared against Gaussian perturbation at matched diversity (σ=0.25), SceneMosaic achieves semantic position score 84.5 vs. 71.6 and relation preservation 99.1% vs. 71.5%, proving diversity does not sacrifice layout validity.
User Study
48 participants rated SceneMosaic highest on both semantic plausibility (4.33) and physical plausibility (4.47), with the lowest standard deviation.
Limitations and Future Work
SceneMosaic relies on single-image initialization, inheriting image-to-3D limitations: one image typically captures only one room, so the method currently operates at room scale. Extending to coherent whole-house environments is a valuable direction. The team believes the division of labor — image prior for speed, agent evolution for accuracy, local decomposition for diversity — is a promising path for large-scale simulation construction. SceneMosaic is open-sourced for community discussion and improvement.
Paper: SceneMosaic: Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution
Authors: Xingjian Ran*, Xiaoye Mo*, Sihao Liu, Jianyu Zhang, Li Luo, Bo Dai
ArXiv: https://arxiv.org/abs/2609.05594
Project: https://rxjfighting.github.io/SceneMosaic
Code: https://github.com/rxjfighting/SceneMosaic
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
