Sparse‑Up: Sparse‑Voxel 3D Geometry Generation Boosts Product Reconstruction

The paper presents Sparse‑Up, a sparse‑voxel based 3D geometry generation framework that introduces learnable surface‑anchored upsampling and view‑partition rendering to eliminate redundant points, dramatically reduce memory usage, and achieve high‑fidelity product reconstruction, with experimental results surpassing TRELLIS, Hunyuan3D and other SOTA methods.

JD Retail Technology
JD Retail Technology
JD Retail Technology
Sparse‑Up: Sparse‑Voxel 3D Geometry Generation Boosts Product Reconstruction

Background

AI‑generated 3D content has progressed using NeRF, Gaussian Splatting, and diffusion models, enabling rapid creation of high‑quality, editable 3D models from simple inputs.

Technical Landscape

Two main architectures dominate AI‑3D generation:

Voxel‑based sparse representation (e.g., Microsoft TRELLIS). Uses sparse voxels driven by 2‑D normals and depth maps. Captures fine details but incurs high compute and memory at high resolution and can produce geometric artifacts.

Point‑cloud / VecSet‑based representation (e.g., Tencent Hunyuan3D). Supervises with SDF values, scales well with data and compute, but struggles with fine‑grained detail reconstruction.

Problem

Analysis of e‑commerce product datasets shows that roughly 70 % of points generated by conventional voxel up‑sampling lie off the true surface, limiting resolution improvement and inflating memory usage.

Proposed Method: Sparse‑Up

Sparse‑Up introduces two innovations to address the bottleneck:

Learnable surface‑anchored upsampling : Replaces octree splitting with a learnable mapping that constrains upsampled points onto the true surface. A Transformer‑based mask generator predicts per‑voxel masks, removing about 70 % of off‑surface points. Integrated with a Texture‑VAE encoder‑decoder, this reduces redundant voxels while improving geometric fidelity.

View‑partition rendering : Divides the input image into overlapping blocks and restricts rendering and back‑propagation to voxels inside the current view frustum. This localized gradient update dramatically lowers GPU memory consumption without sacrificing rendering quality.

The end‑to‑end pipeline takes a single image and directly outputs a textured mesh. The architecture is illustrated in Fig 3.

Sparse‑Up architecture diagram
Sparse‑Up architecture diagram

Experiments

Datasets and Baselines

Training data: 120 k high‑quality samples from Objavarse‑XL, ABO, 3D‑FUTURE and JD’s own furniture/electronics collection. Test sets: Toys4K and a self‑built e‑commerce benchmark. Baselines: TRELLIS, Hunyuan3D 2.1, Step1X‑3D and other open‑source SOTA methods.

Metrics

Reconstruction accuracy: PSNR, LPIPS, SSIM. Generative distribution quality: FID, KID.

VAE Reconstruction (Table 1)

Increasing voxel resolution from 64 to 512 improves PSNR, LPIPS and SSIM consistently; the 512‑resolution model achieves the best texture fidelity.

End‑to‑End Generation (Table 2)

Sparse‑Up outperforms all baselines on FID and KID, confirming superior generative quality.

Memory Ablation (Table 3)

On an 80 GB H100 GPU, surface‑anchored upsampling and view‑partition rendering each reduce memory usage substantially, enabling high‑resolution training.

Qualitative Results

Figures 4‑6 demonstrate that Sparse‑Up preserves fine texture details and yields meshes closely matching real images, even for complex products such as furniture with intricate structures.

Deployment

The algorithm has been integrated into JD’s “MiaoDa” XR product, covering categories such as home furniture, appliances and industrial items, providing a low‑cost, high‑efficiency pipeline for 3D content creation.

Conclusion

Sparse‑Up shows that learnable surface‑anchored upsampling combined with view‑partition rendering can break voxel resolution limits while keeping memory consumption low. Experiments demonstrate significant gains in product geometry reconstruction and overall generative quality. Future work includes expanding industry‑scale 3D datasets, designing next‑generation generation architectures and broadening product‑category coverage.

Reference

Sparse‑Up: Learnable Sparse Upsampling for 3D Generation with High‑Fidelity Textures , ICASSP 2026.

https://arxiv.org/abs/2509.23646
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Neural RenderingSparse VoxelICASSP 20263D Geometry GenerationAI3DProduct Reconstruction
JD Retail Technology
Written by

JD Retail Technology

Official platform of JD Retail Technology, delivering insightful R&D news and a deep look into the lives and work of technologists.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.