Riemannian Deep Learning: Modules, Networks, and Geometry

The article reviews Ziheng Chen’s PhD thesis on Riemannian Deep Learning, outlining a three‑layer framework that unifies manifold‑aware modules, geometry‑specific network designs, and learnable Riemannian metrics, and discusses theoretical foundations, batch normalization, classification heads, specialized networks, and extensive experiments across vision, signal, and graph domains.

Data Party THU
Data Party THU
Data Party THU
Riemannian Deep Learning: Modules, Networks, and Geometry

Introduction

The dissertation Riemannian Deep Learning: Modules, Networks, and Geometries formulates a systematic framework for transferring core deep‑learning components—normalization, fully‑connected layers, convolutions, classification heads, residual structures, and optimizers—onto Riemannian manifolds while preserving interpretability, computational tractability, and training stability.

Why Riemannian Geometry?

Many data types naturally reside on non‑Euclidean spaces: symmetric positive‑definite (SPD) matrices for covariance, rotation matrices on SO(n), subspaces on Grassmann manifolds, hyperbolic spaces for hierarchical data, and correlation matrices for compact statistical representations. Flattening these objects into Euclidean vectors can break positive‑definiteness, unit‑norm constraints, or curvature structure and may introduce approximation errors and numerical instability.

Three‑Layer Contribution Framework

The work organizes contributions into three interconnected layers:

Cross‑manifold unified modules : Generic components that can be instantiated on multiple manifolds (e.g., Lie Group Batch Normalization, Gyrogroup Batch Normalization, Riemannian Multinomial Logistic Regression).

Manifold‑specific network designs : Bespoke layers that exploit the particular geometry of a manifold when generic constructions are insufficient (e.g., Proper Velocity Neural Networks for hyperbolic space, Hyperbolic Busemann Neural Networks, full‑rank correlation matrix networks).

Learnable underlying geometry : Construction or learning of Riemannian metrics better suited for deep networks (e.g., Adaptive Log‑Euclidean Metrics, Product Cholesky Geometries, Bures‑Wasserstein‑Cholesky Metrics).

These layers form a closed loop: migrate Euclidean modules to manifolds, build specialized networks, then adapt the geometry itself.

Mathematical Foundations (Chapter 2)

Chapter 2 provides the toolbox needed for Riemannian deep learning, covering topology, differential geometry, Riemannian and metric geometry, algebraic structures on manifolds, Riemannian optimization, and matrix functions. It answers three practical questions:

Where do data objects reside? (SPD manifold, full‑rank correlation manifold, Grassmann, SO(n), constant‑curvature spaces, etc.)

How are computations performed on these spaces? (Geodesics, exponential/log maps, Fréchet means, group or gyrovector operations replace Euclidean addition, subtraction, scaling.)

How is back‑propagation carried out? (Differentials of matrix functions, Riemannian gradients, tangent‑space mappings, and manifold‑aware optimization.)

Riemannian Batch Normalization (Chapter 3)

Standard batch normalization controls mean and variance in Euclidean space. Directly applying Euclidean formulas on manifolds either destroys structure (if approximated in tangent space) or requires manifold‑specific derivations. The thesis first proposes Lie Group Batch Normalization , which uses group translation for centering and scaling in the tangent space at the identity, preserving the manifold’s geometry. Concrete instances are given for SPD, rotation, and correlation matrices, together with theoretical control of Riemannian sample mean and variance.

Because many important manifolds lack a natural Lie‑group structure, the work introduces pseudo‑reductive gyrogroups , unifying classic gyrogroups and Lie groups. Based on this, Gyrogroup Batch Normalization (GyroBN) generalizes “center‑bias‑scale” to curvature spaces such as Grassmann, constant‑curvature, and correlation manifolds.

LieBN and GyroBN diagrams
LieBN and GyroBN diagrams

Riemannian Classification Layer (Chapter 4)

Euclidean multinomial logistic regression (MLR) is a fully‑connected layer plus softmax, with hyperplane decision boundaries. On manifolds, points and decision boundaries reside in curved spaces, requiring a new definition of “margin”. The thesis first derives an SPD‑specific classifier using geodesic margins and obtains closed‑form expressions under several pull‑back metrics, providing a geometric interpretation of LogEig MLR.

It then generalizes to arbitrary Riemannian manifolds by rewriting the classification formula via Riemannian triangle geometry, resulting in Riemannian Multinomial Logistic Regression (RMLR) that only needs the logarithmic map. This unifies many existing manifold classifiers and creates new ones for SPD and SO(n).

RMLR illustration
RMLR illustration

Manifold‑Specific Networks (Chapter 5)

When generic constructions cannot fully exploit a manifold’s structure, bespoke layers are introduced. Three families are presented:

Proper Velocity Neural Networks (PVNN) : Provide an unconstrained representation for hyperbolic space, enabling stable parameterization of points and construction of MLR, fully‑connected, convolution, activation, and normalization layers.

Hyperbolic Busemann Neural Networks (HBNN) : Use Busemann functions and horospheres to define logits as distances to horospheres, yielding more faithful hyperbolic classifiers with good computational efficiency.

Full‑rank Correlation Networks (CorNet) : Treat correlation matrices as first‑class representations, deriving accurate Riemannian back‑propagation for MLR, fully‑connected, and convolution layers.

PVNN, HBNN, CorNet diagrams
PVNN, HBNN, CorNet diagrams

Designing Learnable Geometry (Chapter 6)

When network modules depend on a Riemannian metric, the metric itself can become a learnable component.

Adaptive Log‑Euclidean Metrics (ALEM) : Parameterize the matrix logarithm within a pull‑back framework, allowing the metric to adapt to data and training dynamics while preserving closed‑form Riemannian operators and an Abelian Lie‑group structure.

Product Cholesky Geometries : Exploit the product structure of Cholesky factors to build Power‑Cholesky and Bures‑Wasserstein‑Cholesky metrics, avoiding costly scalar logarithms and exponentials, yielding fast, stable, closed‑form Riemannian and gyro operators used in SPD classification, residual learning, and tensor interpolation.

ALEM and Cholesky geometry diagrams
ALEM and Cholesky geometry diagrams

Experimental Coverage

Experiments span vision, action recognition, radar, EEG, graph learning, image classification, and genomic sequence modeling. Key findings include:

LieBN and GyroBN improve training stability and performance across diverse backbones compared with naive tangent‑space approximations.

LogEig, SPD MLR, and RMLR achieve higher accuracy on SPD‑based architectures (SPDNet, RResNet) while preserving geometric consistency.

PVNN and HBNN boost hierarchical modeling in graph and genomic tasks; CorNet demonstrates that correlation matrices can serve as effective first‑class inputs.

Adaptive and product‑Cholesky metrics reduce computational overhead, lower failure probability, and mitigate swelling effects in geodesic interpolation.

Conclusion and Future Directions

The thesis consolidates three layers of contribution: (1) unified cross‑manifold modules (LieBN, GyroBN, RMLR); (2) geometry‑specific networks (PVNN, HBNN, CorNet); (3) learnable underlying metrics (ALEM, Product Cholesky, Bures‑Wasserstein‑Cholesky). Future work includes developing mixed‑curvature and matrix‑manifold representations for complex relational structures and integrating Riemannian geometry with generative models such as VAEs, normalizing flows, score‑based models, and optimal transport.

Code example

来源:专知
本文
约5000字
,建议阅读
8
分钟
本文
把黎曼深度学习分成三个互相连接的层次。
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

geometric deep learningbatch normalizationhyperbolic networksmanifold neural networksRiemannian deep learningRiemannian metrics
Data Party THU
Written by

Data Party THU

Official platform of Tsinghua Big Data Research Center, sharing the team's latest research, teaching updates, and big data news.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.