AlphaSchema: Semantic Space Exploration Reshapes LLM Alpha Mining
AlphaSchema introduces a structured trading semantic space with five-dimensional schema plans to decouple exploration from LLM implementation, using a surrogate model and quota-based selection to discover high-performing factors that outperform baselines on China A-share markets while demonstrating robustness across LLM backends.
Background and Problem
Alpha mining aims to discover trading signals with predictive power, executability, interpretability, and out-of-sample stability. Traditional asset pricing derives factors from economic hypotheses, but the space of latent trading mechanisms is vast and evolving, limiting manual discovery. Automated systems have evolved through formula-based evolutionary search — which explores implementation spaces via mathematical combinations but struggles to inject high-level trading hypotheses — and LLM-based agent systems that couple search-space construction, trajectory selection, and code implementation inside the LLM, making diversity hard to measure, coverage hard to control, and compute costs high.
AlphaSchema's core insight is to change the search abstraction level: define and explore trading hypotheses at the semantic level first, then let an LLM translate selected semantic plans into executable factors. This makes the search space explicit, measurable, and optimizable while retaining robustness to the choice of LLM.
Method
4.1 Structured Semantic Plan Space
Each candidate factor is represented as a 5-tuple p = (e, c, Q, d, o) ∈ P:
Event (e) : market phenomenon or signal-generating event (e.g., breakout, volume surge, volatility compression).
Context (c) : market state or reference condition for interpreting the event (e.g., recent extremes, VWAP region, volatility regime).
Qualities (Q) : additional attributes refining the mechanism (e.g., volume confirmation, multi-period consistency, outlier filtering); 0–3 per plan.
Direction (d) : expected relationship with future returns (e.g., continuation, reversal, range-bound).
Output (o) : form in which the mechanism is expressed as a tradable signal (e.g., continuous score, event decay, cross-sectional rank).
The plan space is the Cartesian product P = V_E × V_C × Q_set × V_D × V_O where Q_set = {Q ⊆ V_Q : 0 ≤ |Q| ≤ 3}. In experiments the price-volume semantic space contains 140 components : 40 events, 40 contexts, 50 qualities, 3 directions, and 7 outputs. A plan only describes trading logic; it does not prescribe operators, library functions, or numeric parameters. No hand-crafted compatibility rules are imposed — unusual or potentially inconsistent combinations are evaluated through code implementation and empirical feedback, preserving unconventional combinations that may yield useful factors.
4.2 Plan Implementation and Execution Guard
Given a selected plan, a code agent receives its semantic specification, data contract, and implementation constraints and translates it into an executable factor function. Each plan is implemented at two time scales (fast and slow), producing f_{p,fast} and f_{p,slow}. Before backtesting, every implementation must pass an execution guard that verifies contract compliance, tests numerical stability, and screens for look-ahead leakage. Invalid implementations get one repair attempt; if they still fail they receive zero reward — treating irrecoverable implementation failure as negative semantic feedback.
4.3 Semantic Reward Modeling
For each valid implementation f_{p,s} of plan p, the reward is defined as:
r_{p,s} = α · RankIC(f_{p,s}) + β · RankICIR(f_{p,s}) − λ · Δ_{lag}(f_{p,s})where Δ_{lag}(f) = max{0, RankIC(f) − RankIC(L_1 f)} measures signal lag sensitivity. Coefficients are set to (α, β, λ) = (10, 1, 2), controlling predictive strength, temporal stability, and lag sensitivity respectively. The plan reward takes the maximum over the fast and slow implementations.
After each search round, new plan-reward pairs are added to a cumulative buffer . A LightGBM reward model is retrained on the buffer using structured plan features φ(p) — including one-hot encodings, category counts, and pairwise component interactions. All historical observations are retained (no sliding window) because each plan-level supervision is expensive and old labels remain valid under a fixed evaluation protocol.
4.4 Adaptive Quota Selection Mechanism
Each round's budget B is allocated among three strategies. The exploration fraction decays with the number of evaluated plans: ρ_t = ρ_min + (ρ_max − ρ_min) · exp(−n_t / τ) After a cold-start phase (full budget to exploration), ρ_t gradually decreases from ρ_max to ρ_min; the remaining budget is split between exploitation and mutation.
Exploration : prefers structural novelty, prioritizing under-covered event-context pairs. Exploration score combines evaluated counts of event-context pairs, individual events, and individual contexts.
Surrogate-guided exploitation : ranks candidate pool by surrogate model predictions.
Local mutation : starting from high-reward plans, changes one schema component to generate unseen neighbors — replace quality (0.22), replace context (0.16), replace output (0.16), add quality (0.14), remove quality (0.12), replace direction (0.12), replace event (0.08).
Each round samples 10,000 candidate plans from the semantic space, deduplicates, and allocates by quota. The final factor pool is built by a greedy filter on reward-sorted factors: a factor is added only if its absolute correlation with every factor already in the pool is below 0.7. Downstream, a LightGBM ranker ensembles the factor pool using a Top50/Drop5 trading strategy.
Experiments
Experimental Setup
Dataset : CSI300 universe; training 2016-01-01 to 2020-12-31, validation 2021-01-01 to 2022-12-31, test 2023-01-01 to 2025-12-31. Prediction target: 5-day forward close-to-close return.
Search configuration : price-volume semantic space (140 components). Default code agent: DeepSeek-V4-Flash. 16 plans evaluated per round, 80 rounds total, first 10 rounds pure exploration. 5 independent runs.
Baselines : ML predictors (MLP, XGBoost), deep sequence models (Transformer, GRU, LSTM), open-source factor libraries (Alpha158, Alpha360), agent mining systems (RD-Agent, QuantaAlpha).
Main Results
AlphaSchema (OHLCV) with 120 discovered factors achieves the best IC = 0.0382 and ICIR = 0.2374 among all methods. The +Fundamental variant adding 30 fundamental factors (150 total) further improves portfolio performance, achieving the best IR = 1.0877 and AER = 11.94% .
Key comparisons:
vs. LSTM (best Rank IC baseline): AlphaSchema IC +0.4%, ICIR +4.6%, IR +72%.
vs. RD-Agent (best IR among agent systems): AlphaSchema(+Fund.) IR +10.3%, AER +75.3%.
vs. Alpha158 (best IC among open-source libraries): AlphaSchema IC +10.1%, ICIR +14.1%, IR +98.7%.
NAV trajectories show AlphaSchema's curve maintains higher net-value growth throughout the test window.
Analysis Experiments
Semantic Component Ablation
Leave-one-out ablation on 100 full schema plans. Removing any semantic component degrades factor quality, while implementation success rate drops only moderately. Retained signal strength relative to full plan: full 100%, minus event 78.9%, minus quality 72.9%, minus direction 71.7%, minus context 69.4%, minus output 75.1%. This shows the five dimensions provide complementary information — the LLM can compensate operational details from remaining fields but cannot fully recover missing semantic constraints.
Semantic Navigation Analysis
All plans visited during search are embedded into a 50-dimensional PCA space and K-means clustered into 12 regions. Search trajectories show: early stages cover diverse semantic regions (compression, local extremum, breakout-type events); later stages progressively reallocate evaluation budget to high-reward regions (short-long efficiency conflict, multi-window alignment patterns). Average reward rises from 1.415 to 2.516. Search does not collapse to a single narrow cluster but maintains multiple active exploration regions.
Implementation Budget Efficiency
For 100 plans, up to 8 implementations each. Single-implementation Top-20% recall is 0.462; 5 implementations raise recall to 0.780, but per-implementation recall efficiency drops from 0.462 to 0.156. Conclusion: single implementation is the budget-efficient choice for large-scale alpha mining — the surrogate model can aggregate noisy single-shot rewards at the feature level.
LLM Implementation Robustness
Fixed 100 schema plans, implemented with 7 different LLM backends. Pass@1 varies significantly across models, but the mean |RankIC| among successful implementations stays in a narrow band of 0.0116–0.0168 with no monotonic relationship to model capability. Stronger LLMs mainly improve implementation success rate; factor quality is primarily determined by the schema plan being implemented.
Factor Decay Analysis
Over the 2023–2025 test period, AlphaSchema's 150-factor pool maintains higher Rank IC than Alpha158 on 466 out of 467 valid rolling dates, with an average gap of 0.0149.
Semantic Space Reward Predictability
On 12,126 valid plan-reward pairs, LightGBM predicting reward from structured schema features achieves Spearman correlation 0.239 and Top10% lift 1.519, while a label-shuffled control approaches random (0.029), confirming learnable reward structure exists in the semantic space.
Key Conclusions
AlphaSchema demonstrates that decoupling semantic exploration from LLM implementation yields a controllable, measurable, and optimizable search process for alpha mining. The structured semantic plan space (Event, Context, Qualities, Direction, Output) provides complementary constraints that LLM implementation alone cannot recover. A LightGBM surrogate model trained on cumulative plan-reward observations enables efficient navigation via an adaptive quota mechanism balancing exploration, exploitation, and mutation. Empirical results on China A-share data show state-of-the-art predictive and portfolio metrics, robustness across LLM backends, and sustained outperformance over existing factor libraries and agent systems. The framework establishes that alpha mining quality is determined by the semantic plan, not the specific LLM used for implementation.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
