MAPLE: Single-Training Multi-Alpha Framework for Diversified Portfolios
MAPLE introduces a backbone-agnostic framework that generates multiple diverse alpha signals in a single training run by combining a unified capacity-scaled prediction head, extreme rank-weighted listwise loss, and explicit diversity regularization, achieving superior Sharpe and Calmar ratios across four global equity markets with 55x fewer parameters than MoE baselines.
Abstract
Classic alpha mining combines many low-correlation predictive signals to achieve strong risk-adjusted returns, but deep learning stock-ranking methods typically produce only one alpha per stock. They rely on increasingly complex architectures with diminishing returns and can only obtain diversity by training multiple independent models or using implicit routing, without explicit control over inter-alpha correlation.
This paper proposes MAPLE (Multi-Alpha Position-aware Listwise Ensembling) , a backbone-agnostic framework that recovers the classic diversity principle in a single training run. MAPLE combines a unified capacity-scaled prediction head with an extreme rank-weighted listwise ranking loss and introduces a diversity regularizer that explicitly penalizes pairwise alpha correlations. Across four equity markets (US, China, Japan), MAPLE achieves the best average Sharpe ratio (1.690) and Calmar ratio (2.175) among nine baselines, using 55x fewer parameters (186K vs. 10.3M for DHMoE) and training 2.5x faster. It generalizes across five backbone architectures, improving Sharpe by 10–23% and Calmar by 17–43%. Behavioral analysis reveals that the unified prediction head reduces inter-alpha correlation before any diversity loss is applied; the extreme rank loss enables diversity regularization to improve rather than erode single-alpha ranking quality; and capacity scaling maintains this balance as the model grows. These results show that principled loss design and capacity allocation—not architectural complexity—drive diverse and effective multi-alpha generation.
Background
Classic alpha mining aims to discover predictive signals that can be turned into trading strategies and combined into portfolios. To overcome manual discovery limits, researchers have explored genetic programming, reinforcement learning, and LLM-based agents. Deep learning stock ranking has become the dominant paradigm: models predict cross-sectional returns for each stock and build portfolios by selecting top-ranked stocks.
Two fundamental limitations remain unresolved. First, diminishing returns from architectural complexity. Recent work stacks multiple Transformer layers with inter-stock attention, or adds graph networks, MoE routing, and diffusion components, greatly increasing compute for marginal gains. Simpler architectures have been shown to match or exceed complex Transformer variants on financial tasks, and these architecture-specific designs do not transfer across backbones. Second, neglect of signal diversity. The core principle of classic alpha mining is signal diversity: single alphas are vulnerable to regime shifts, and low-correlation ensembles are crucial for managing concentrated exposure. Yet most deep learning methods produce only one alpha score per stock. Existing remedies—training multiple independent models, explicit routing mechanisms, or multi-stage distillation—multiply training cost and do not explicitly control inter-alpha correlation. This raises the question: can we simultaneously achieve single-model prediction efficiency and diversified multi-alpha generation without complex architectures or multi-stage pipelines?
Problem Formulation
Portfolio construction is cast as a learning-to-rank task with future returns as targets. Given historical information X ∈ R^{S×T×F} (S stocks, T time steps, F features), model f_θ learns to predict q-day future returns y ∈ R^S: min_θ ℒ(y, f_θ(X)) At inference, the top-k predicted stocks each day form a long portfolio. The model f_θ comprises three sequential components: a temporal encoder t_θ extracting per-stock representations H ∈ R^{S×D}; a cross-sectional module g_θ capturing inter-stock relations H' = g_θ(H); and a prediction head h_θ producing alpha scores ŷ = h_θ(H') ∈ R^S. In existing work h_θ is usually a simple linear projection yielding a single ranking score per stock. MAPLE redesigns g_θ and h_θ into a unified module that produces diversified alpha signals in a single forward pass.
Method
MAPLE consists of three core components: a unified multi-alpha generation head, an extreme rank-weighted position-aware listwise ranking loss, and a diversity regularizer, complemented by a capacity scaling scheme.
4.1 Multi-Alpha Generation Head
Given per-stock hidden representations H ∈ R^{S×D} from any backbone, MAPLE jointly projects them into N_α alpha signals in one forward pass. The hidden representations are first LayerNorm-normalized H̃ = LayerNorm(H), serving as shared input to two paths that are summed in the prediction space .
Intra-Stock Path: H̃ passes through a two-layer MLP (ReLU activation) projecting to N_α predictions, with hidden dimension set to ⌊D·(N_α/8)⌋, scaling proportionally with the number of alphas: Ŷ_intra = MLP(H̃) ∈ R^{S×N_α}.
Inter-Stock Path: Multi-head attention is computed along the stock axis, assigning one attention head per alpha ( h = N_α), giving each alpha independent query and key parameters. The output of head i is: A^{(i)} = softmax(Q^{(i)} K^{(i)⊤} / √d_k) · V^{(i)}. Each head independently produces a scalar ranking score per stock, concatenated to form Ŷ_inter ∈ R^{S×N_α}.
Prediction-Level Residual Aggregation: Unlike standard Transformer blocks that apply residual connections in hidden space, the two paths sum in output space: Ŷ = Ŷ_intra + Ŷ_inter ∈ R^{S×N_α}. This design shifts path contributions from overlapping to near-orthogonal—standard residual path contribution coefficients sum to 1.136, while this design sums to 1.008—reducing inter-alpha correlation from 0.998 to 0.944 before any diversity loss is applied, providing a favorable diversity starting point.
4.2 Trainable Position-Aware Listwise Ranking
Portfolio construction relies on relative ranking of assets rather than exact return values, so a listwise objective directly maximizes Spearman rank correlation between predictions and targets across all S stocks:
ℒ_spearman = -(1/N_α) · Σ_{i=1}^{N_α} ρ(φ(ŷ_i), φ(y))However, the global Spearman objective conflicts with diversity regularization. To resolve this, MAPLE introduces the extreme rank-weighted Spearman loss , which concentrates each alpha's learning signal on extreme rank positions, allowing diversity regularization to differentiate alphas only within a smaller profitable region. Each alpha i has two learnable parameters: sharpness ξ^{(i)} controlling transition steepness, and scale γ^{(i)} determining margin δ^{(i)} = γ^{(i)}·(S-1)/2. Weights approach 1 at top ranks and 0 at center and below:
ℒ_extreme = -(1/N_α) · Σ_{i=1}^{N_α} C^{(i)} · Σ_{s=1}^{S} φ̃(ŷ_i)_s · φ̃(y)_s · v_s^{(i)}where C^{(i)} = 1 + 2δ^{(i)}/S is a coverage ratio compensating for reduced gradient magnitude when focusing on extreme positions. ℒ_extreme concentrates learning signals without forcing ranking at other positions, so ℒ_spearman is retained to maintain global ranking quality.
4.3 Diversity-Controlled Multi-Alpha Ensembling
Diversity Regularization penalizes redundancy among alphas by discouraging pairwise absolute correlation of predictive signals:
ℒ_diversity = (1/(N_α(N_α-1))) · Σ_{i≠j} |ρ(φ(ŷ_i), φ(ŷ_j))|Absolute value is used because two strongly negatively correlated alphas would select opposite ends of the ranking under long-only constraints, making one a poor alpha rather than a diversifier. The full training objective combines both ranking losses and the diversity regularizer:
ℒ = ℒ_spearman + ℒ_extreme + λ · ℒ_diversitywhere λ controls the trade-off between single-alpha predictive quality and inter-alpha diversity (default λ = 0.1). Alpha Aggregation: At inference, each alpha independently selects the top-k stocks to form a portfolio; the final return is the equal-weighted average of all N_α portfolios, consistent with classic alpha mining practice.
Experiments
5.1 Experimental Setup
Datasets: Evaluated on four major equity markets—CSI300 and CSI500 (China, 295 and 514 stocks), NI225 (Japan, 209 stocks), SP500 (US, 525 stocks). Eight time-series features (OHLCV and moving averages) are used. Training period 2008–2019, validation 2020, test 2021–2024.
Baselines: Nine baselines covering diverse architecture families: RankLSTM (intra-stock only); FinGAT, MASTER, StockMixer, CI-STHPAN (graph, attention, MLP-mixer, hypergraph inter-stock mechanisms); MERA and DHMoE (MoE routing, DHMoE adds diffusion module); AlphaMix (explicit multi-model ensemble); TIPS (multi-architecture distillation). Metrics: Annualized Return (AR), Sharpe Ratio (SR), Calmar Ratio (CR).
5.2 Main Results
MAPLE achieves the best average portfolio performance across four markets with average SR=1.690, CR=2.175, surpassing all baselines—including the latest MoE methods MERA, DHMoE, and distillation method TIPS—and attains the highest SR in each individual market. Per-market results: CSI300 SR=1.851, CR=2.062; CSI500 SR=2.161, CR=1.961; NI225 SR=0.991, CR=0.783; SP500 SR=1.758, CR=3.896.
Efficiency advantages are pronounced: MAPLE uses only 186K parameters—about 55x fewer than DHMoE (10.3M) and nearly 8x fewer than MERA (1.46M). Per-epoch training time is 3.60 seconds, roughly 2–2.5x faster than the three heaviest baselines. Inference FLOPs are only 0.130G, one to two orders of magnitude lower than attention and MoE baselines. This confirms that architectural complexity does not translate into proportional performance gains, while lighter baselines lag MAPLE by at least 0.23 SR—a meaningful gap given that portfolio models require frequent retraining.
5.3 Ablation Studies
First group—Lightweight Architecture: Simply increasing alpha outputs from 1 to 8 improves SR from 1.402 to 1.434, confirming the value of multi-alpha diversity. Adding lightweight attention raises SR and CR; adding prediction-level residual aggregation improves all three metrics (SR 1.483, CR 1.826) and outperforms a standard Transformer block (SR 1.305, CR 1.511).
Second group—Diversified Ensembling: Combining naive Spearman with diversity regularization actually degrades performance (SR 1.415, CR 1.590), confirming the conflict. Replacing with extreme rank loss improves all metrics (SR 1.510, CR 1.856); adding diversity control further boosts (SR 1.579, CR 1.996). Scaling single-alpha capacity (SR 1.660, CR 2.031) and expanding to 24 alphas yields full MAPLE (SR 1.690, CR 2.175).
5.4 Backbone Generalization
Beyond the default Transformer backbone, MAPLE is attached to TCN, GRU, LSTM, and Mamba encoders. Across all five backbones, average SR and CR improve, with SR gains ranging from 10% (Mamba) to 23% (GRU), and CR gains from 17% (Mamba) to 43% (GRU) . This demonstrates that MAPLE's gains stem from principled design choices rather than architecture-specific tuning, enabling reliable cross-backbone transfer.
5.5 Behavioral Analysis: Why Each Component Works
Better Inter-Stock Attention Structure: Prediction-level residuals bring the two paths' gradient norm ratios closer (1.20 vs. 4.38 for standard Transformer), shifting path contributions from overlapping to near-orthogonal. This improves alpha diversity rather than single-alpha strength—inter-alpha correlation drops from 0.998 to 0.944, yielding real diversification gains (ensemble SR 1.483 vs. 1.434) without explicit diversity optimization.
How Extreme Ranking Enables Diversity: Using extreme rank loss alone makes all alphas converge to the same top stocks, achieving highest precision but almost no diversification. Adding diversity regularization on top of Spearman reduces overlap but by dispersing selections, causing precision to drop (0.223→0.203). Imposing diversity regularization on top of Spearman+extreme rank, precision drops only 5.1% instead of 9.0% , because each alpha's learning signal is already concentrated, so diversity regularization only needs to differentiate alphas within the profitable region, achieving the highest return-to-volatility ratio (0.112).
Scaling with Alpha Count: Without capacity scaling, the loss design fails to benefit from more alphas; SR plateaus after N_α=4 and is surpassed by multi-seed ensembles. With capacity scaling, AR resumes growth after N_α=4, and SR matches or exceeds multi-seed ensembles in most settings—a single trained model matches the risk-adjusted performance of an explicit ensemble, with annualized return roughly double that of the ensemble.
Diminishing Returns of Diversity Weight: λ exhibits an inverted-U shape for CR, peaking at λ∈[0.1, 0.15] and declining monotonically for λ>0.2. The correlation level naturally reached by independent training is what the regularizer should approach but not exceed—exceeding it only sacrifices single-alpha ranking quality without corresponding gains. Capacity scaling and alpha scaling make it easier in practice to find and maintain the λ that reaches the natural diversification region.
Paper link: https://arxiv.org/pdf/2607.24131
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
