How Dewu Delivers 60FPS 3D Spatial Photos on Mobile with Gaussian Splatting
Dewu's engineering team details their end-to-end pipeline for turning single photos into interactive 3D spatial images using 3D Gaussian Splatting, covering model generation, perceptual quality evaluation, 81.5% model compression, and mobile GPU optimizations that achieve stable 60FPS rendering on iOS and Android.
Dewu's audio-video team implemented a complete 3D spatial photo feature where users post a single image tagged #空间图片, the server asynchronously generates a 3D Gaussian scene, and the client renders it in real-time based on device pose — all while preserving the familiar image-posting workflow.
End-to-End Pipeline
The system balances model quality, file size, and client-side cost. The server handles generation, multi-view rendering, quality assessment, and compression; the client performs device qualification, download, decoding, GPU upload, and real-time rendering.
Model Production: Single-Image to 3D Gaussian Scene
Motion Parallax Principle
Static images capture color from a fixed viewpoint. When the observer moves, objects at different distances shift by different amounts — this motion parallax creates depth perception. The system builds a renderable 3D representation so that changing the virtual camera viewpoint produces the correct parallax.
3D Gaussian Representation
3D Gaussian Splatting (3DGS) represents a scene as a set of 3D Gaussian primitives. Each primitive is a soft ellipsoid with position, scale, rotation, color, and opacity. During rendering, Gaussians are projected to 2D screen-space ellipses, sorted by depth, and alpha-blended. Changing the camera alters projection and occlusion, producing parallax between foreground and background. Gaussian count, screen coverage, and overlap directly affect draw cost.
Unlike classic 3DGS which optimizes Gaussians from multi-view photos, this approach predicts Gaussian attributes directly from a single image via a feed-forward network, then hands them to the renderer for nearby viewpoints — no re-generation on every phone movement.
Single-Image Generation Boundaries
Single-image generation lowers the content barrier but lacks ground truth for occluded regions. New viewpoints contain hallucinated content; larger viewpoint changes expose holes, stretching, or wrong occlusion. The experience is therefore constrained to a small neighborhood around the original view. Generation resolution also controls Gaussian count, which cascades into storage, sorting, and draw overhead — generation resolution and client output resolution are tuned independently.
Quality Evaluation: From Pixel Error to Multi-View Perception
Quality must be assessed across viewpoint changes. A shoelace may look intact front-on but break when rotated; silhouettes may show holes; distant textures may become sparse.
Pixel Metrics
PSNR measures mean squared pixel error (numerical fidelity). SSIM compares local luminance, contrast, and structure but still relies on spatial correspondence. Small geometric shifts move edges, hurting pixel metrics even when perception is good; over-smoothed results can score well yet lose detail. Therefore, perceptual metrics are primary.
Perceptual Metrics
LPIPS (Learned Perceptual Image Patch Similarity) uses a neural network to compare deep features, capturing texture and contour changes. DISTS (Deep Image Structure and Texture Similarity) combines structural and texture statistics from deep features, tolerating mild resampling and geometric variation. Both are lower-is-better. Reference images must be meaningful: comparing a novel-view render directly against the original photo would count correct parallax as error.
Multi-View Consistency Testing
Two reference types are used: (1) original photo for generation quality, (2) baseline model renders for compression quality. Comparisons require aligned resolution and crop. Frame-level metrics don't capture temporal flicker; continuous viewing and inspection of occlusion edges, text, and fine structures are also needed.
Model Compression: Balancing Quality, Size, and Render Cost
Gaussian Reduction vs. Quality
Lightweighting has two stages: reduce primitive count, then compress per-primitive storage. Reducing points affects both file size and client workload; format encoding mainly shrinks distribution size, with runtime gains depending on decode path and data layout.
Experiments compared lowering generation resolution, uniform downsampling, and filtering primitives. Uniform downsampling caused visible striping when zoomed. Lowering generation resolution lost detail in hair, silhouettes, and large backgrounds at very low counts. Evaluation used a fixed 5-second render video (1 fps → 185 frame pairs per configuration) with LPIPS and DISTS. Higher Gaussian counts consistently achieved higher frame-level pass rates on both metrics. Visual review confirmed that snow scenes and distant neighborhoods exhibited black shadows, holes, or missing elements at low counts, which improved at the chosen higher count.
Attribute Encoding & Format Compression
After fixing model scale, the raw PLY is compressed via attribute quantization and compact encoding. In the tested layout, each Gaussian carries 14 float32 values (position, color, opacity, scale, rotation) ≈ 56 bytes. Optimized encoding reduced file size by 81.5% versus the uncompressed baseline while preserving render quality.
Mobile Rendering: Fluid Response to Viewpoint Changes
Thousands of Gaussians on mobile incur cost in decoding, data transfer, sorting, and pixel blending. The pipeline parallelizes attribute decode, builds a depth index on first frame, then reuses it for subsequent draws.
Coordinate Space & View Interaction
Gaussians live in world space. An initial camera is built from model extrinsics and a look-at target; intrinsics and aspect define projection. The view matrix V transforms Gaussians to camera space (forward = -Z), then projection matrix P maps to screen. Gyroscope integrates angular velocity into constrained lateral/vertical camera offsets, always looking at the target, creating small parallax. Gestures use orbit/pan/zoom for larger 3D navigation.
GPU Depth Sorting
Projected Gaussians are semi-transparent ellipses; multiple can cover one pixel. Alpha blending is order-dependent: far-to-near drawing is required. The system sorts by Gaussian center depth in camera space — a global approximation, not per-pixel exact occlusion. A GPU stable Radix Sort generates 32-bit depth keys, processes 8 bits per pass (4 passes: count, prefix sum, rearrange), and outputs a sorted index buffer that stays on GPU. Compute-to-render resource dependencies are established. The index is reused for small parallax; large viewpoint changes trigger re-sort. GPU sorting cut sort time by 60–75% .
GPU Instanced Drawing
Each Gaussian is a 4-vertex quad. The fragment shader computes elliptical color/opacity. The standard path uses a single instanced draw: vertexCount = 4, instanceCount = N, eliminating per-Gaussian CPU draw calls and vertex expansion. The vertex shader reads gaussians[sortedIndices[instanceID]] via instanceID; vertexID generates the four corners. The same ~4N vertex workload and blending remain, but CPU submission overhead vanishes.
// C++-style pseudocode: single instanced draw for standard path
bind(gaussianBuffer, sortedIndices, uniforms);
drawInstanced(TriangleStrip, 4, pointCount);
// Vertex stage reads gaussians[sortedIndices[instanceID]].Other Optimizations
Additional optimizations include frustum culling, early-z testing, and memory layout tuning (details in article diagrams).
On-Device Measured Performance
iOS and Android high/mid/low-tier devices were tested across Gaussian counts for visual quality, frame rate, memory, and thermals. The higher-count model was selected. Real-device tests show stable 60 FPS on both platforms inside the Dewu app.
Summary
By jointly evaluating multi-view quality and on-device metrics, the team locked a model tier, applied format encoding for an 81.5% distribution size reduction, and used client-side data organization, GPU sorting, and frame scheduling to cut runtime overhead. The result preserves detail and multi-view feel while making 3D content lightweight and fluid enough for everyday browsing and interaction.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
DeWu Technology
A platform for sharing and discussing tech knowledge, guiding you toward the cloud of technology.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
