Mobile Development 17 min read

How Dewu Delivers 60FPS 3D Spatial Photos on Mobile with Gaussian Splatting

Dewu's engineering team details their end-to-end pipeline for turning single photos into interactive 3D spatial images using 3D Gaussian Splatting, covering model generation, perceptual quality evaluation, 81.5% model compression, and mobile GPU optimizations that achieve stable 60FPS rendering on iOS and Android.

DeWu Technology
DeWu Technology
DeWu Technology
How Dewu Delivers 60FPS 3D Spatial Photos on Mobile with Gaussian Splatting

Dewu's audio-video team implemented a complete 3D spatial photo feature where users post a single image tagged #空间图片, the server asynchronously generates a 3D Gaussian scene, and the client renders it in real-time based on device pose — all while preserving the familiar image-posting workflow.

End-to-End Pipeline

The system balances model quality, file size, and client-side cost. The server handles generation, multi-view rendering, quality assessment, and compression; the client performs device qualification, download, decoding, GPU upload, and real-time rendering.

Model Production: Single-Image to 3D Gaussian Scene

Motion Parallax Principle

Static images capture color from a fixed viewpoint. When the observer moves, objects at different distances shift by different amounts — this motion parallax creates depth perception. The system builds a renderable 3D representation so that changing the virtual camera viewpoint produces the correct parallax.

3D Gaussian Representation

3D Gaussian Splatting (3DGS) represents a scene as a set of 3D Gaussian primitives. Each primitive is a soft ellipsoid with position, scale, rotation, color, and opacity. During rendering, Gaussians are projected to 2D screen-space ellipses, sorted by depth, and alpha-blended. Changing the camera alters projection and occlusion, producing parallax between foreground and background. Gaussian count, screen coverage, and overlap directly affect draw cost.

Unlike classic 3DGS which optimizes Gaussians from multi-view photos, this approach predicts Gaussian attributes directly from a single image via a feed-forward network, then hands them to the renderer for nearby viewpoints — no re-generation on every phone movement.

Single-Image Generation Boundaries

Single-image generation lowers the content barrier but lacks ground truth for occluded regions. New viewpoints contain hallucinated content; larger viewpoint changes expose holes, stretching, or wrong occlusion. The experience is therefore constrained to a small neighborhood around the original view. Generation resolution also controls Gaussian count, which cascades into storage, sorting, and draw overhead — generation resolution and client output resolution are tuned independently.

Quality Evaluation: From Pixel Error to Multi-View Perception

Quality must be assessed across viewpoint changes. A shoelace may look intact front-on but break when rotated; silhouettes may show holes; distant textures may become sparse.

Pixel Metrics

PSNR measures mean squared pixel error (numerical fidelity). SSIM compares local luminance, contrast, and structure but still relies on spatial correspondence. Small geometric shifts move edges, hurting pixel metrics even when perception is good; over-smoothed results can score well yet lose detail. Therefore, perceptual metrics are primary.

Perceptual Metrics

LPIPS (Learned Perceptual Image Patch Similarity) uses a neural network to compare deep features, capturing texture and contour changes. DISTS (Deep Image Structure and Texture Similarity) combines structural and texture statistics from deep features, tolerating mild resampling and geometric variation. Both are lower-is-better. Reference images must be meaningful: comparing a novel-view render directly against the original photo would count correct parallax as error.

Multi-View Consistency Testing

Two reference types are used: (1) original photo for generation quality, (2) baseline model renders for compression quality. Comparisons require aligned resolution and crop. Frame-level metrics don't capture temporal flicker; continuous viewing and inspection of occlusion edges, text, and fine structures are also needed.

Model Compression: Balancing Quality, Size, and Render Cost

Gaussian Reduction vs. Quality

Lightweighting has two stages: reduce primitive count, then compress per-primitive storage. Reducing points affects both file size and client workload; format encoding mainly shrinks distribution size, with runtime gains depending on decode path and data layout.

Experiments compared lowering generation resolution, uniform downsampling, and filtering primitives. Uniform downsampling caused visible striping when zoomed. Lowering generation resolution lost detail in hair, silhouettes, and large backgrounds at very low counts. Evaluation used a fixed 5-second render video (1 fps → 185 frame pairs per configuration) with LPIPS and DISTS. Higher Gaussian counts consistently achieved higher frame-level pass rates on both metrics. Visual review confirmed that snow scenes and distant neighborhoods exhibited black shadows, holes, or missing elements at low counts, which improved at the chosen higher count.

Attribute Encoding & Format Compression

After fixing model scale, the raw PLY is compressed via attribute quantization and compact encoding. In the tested layout, each Gaussian carries 14 float32 values (position, color, opacity, scale, rotation) ≈ 56 bytes. Optimized encoding reduced file size by 81.5% versus the uncompressed baseline while preserving render quality.

Mobile Rendering: Fluid Response to Viewpoint Changes

Thousands of Gaussians on mobile incur cost in decoding, data transfer, sorting, and pixel blending. The pipeline parallelizes attribute decode, builds a depth index on first frame, then reuses it for subsequent draws.

Coordinate Space & View Interaction

Gaussians live in world space. An initial camera is built from model extrinsics and a look-at target; intrinsics and aspect define projection. The view matrix V transforms Gaussians to camera space (forward = -Z), then projection matrix P maps to screen. Gyroscope integrates angular velocity into constrained lateral/vertical camera offsets, always looking at the target, creating small parallax. Gestures use orbit/pan/zoom for larger 3D navigation.

GPU Depth Sorting

Projected Gaussians are semi-transparent ellipses; multiple can cover one pixel. Alpha blending is order-dependent: far-to-near drawing is required. The system sorts by Gaussian center depth in camera space — a global approximation, not per-pixel exact occlusion. A GPU stable Radix Sort generates 32-bit depth keys, processes 8 bits per pass (4 passes: count, prefix sum, rearrange), and outputs a sorted index buffer that stays on GPU. Compute-to-render resource dependencies are established. The index is reused for small parallax; large viewpoint changes trigger re-sort. GPU sorting cut sort time by 60–75% .

GPU Instanced Drawing

Each Gaussian is a 4-vertex quad. The fragment shader computes elliptical color/opacity. The standard path uses a single instanced draw: vertexCount = 4, instanceCount = N, eliminating per-Gaussian CPU draw calls and vertex expansion. The vertex shader reads gaussians[sortedIndices[instanceID]] via instanceID; vertexID generates the four corners. The same ~4N vertex workload and blending remain, but CPU submission overhead vanishes.

// C++-style pseudocode: single instanced draw for standard path
bind(gaussianBuffer, sortedIndices, uniforms);
drawInstanced(TriangleStrip, 4, pointCount);
// Vertex stage reads gaussians[sortedIndices[instanceID]].

Other Optimizations

Additional optimizations include frustum culling, early-z testing, and memory layout tuning (details in article diagrams).

On-Device Measured Performance

iOS and Android high/mid/low-tier devices were tested across Gaussian counts for visual quality, frame rate, memory, and thermals. The higher-count model was selected. Real-device tests show stable 60 FPS on both platforms inside the Dewu app.

Summary

By jointly evaluating multi-view quality and on-device metrics, the team locked a model tier, applied format encoding for an 81.5% distribution size reduction, and used client-side data organization, GPU sorting, and frame scheduling to cut runtime overhead. The result preserves detail and multi-view feel while making 3D content lightweight and fluid enough for everyday browsing and interaction.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

model compressionmobile rendering3D Gaussian SplattingDISTSGPU radix sortinstanced drawingLPIPSsingle-image 3D reconstructionspatial photos
DeWu Technology
Written by

DeWu Technology

A platform for sharing and discussing tech knowledge, guiding you toward the cloud of technology.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.