CoLAS: Multimodal Corroboration Mining for Reliable Financial Trading Signals

This paper introduces CoLAS, a framework that explicitly models multimodal corroboration—non-cancelling support across heterogeneous financial modalities—to extract more reliable trading signals, achieving over 25% average annualized return improvement across six datasets including five stocks and Bitcoin.

Bighead's Algorithm Notes
Bighead's Algorithm Notes
Bighead's Algorithm Notes
CoLAS: Multimodal Corroboration Mining for Reliable Financial Trading Signals

Abstract

Financial trading relies on extracting reliable signals from heterogeneous market modalities such as price sequences, breaking news, and investor sentiment. Existing multimodal methods primarily exploit complementarity—information present in one modality but absent in another—while overlooking a critical question: do different modalities provide mutually supporting evidence for the same trading signal? The authors define this "task-conditioned, non-cancelling support" as multimodal corroboration . In finance, this is especially important because individual financial perspectives are noisy and information-weak, while persistent support across heterogeneous views may provide more stable task-relevant signals than evidence appearing in only a single view.

The authors propose the CoLAS framework, which operationalizes multimodal corroboration as a trainable task-conditioned representation. Modality representations are organized into a per-instance matrix; a softmax-based spectral objective reinforces their dominant shared component; signed modality contributions determine whether that component provides non-cancelling support and constructs the corroboration signal; and a coupled robustness-aware consistency objective keeps the corroboration signal stable when modalities are corrupted or missing. Experiments on six datasets (five U.S. stocks and Bitcoin) show CoLAS consistently outperforms existing methods in annualized return and Sharpe ratio.

Background

Data-driven financial trading aims to discover reliable trading signals from various market data. Heterogeneous modalities—price, news, sentiment—characterize the market from different angles, providing complementary signal sources. Prior work has evolved through three stages: (1) Early methods coarsely combine modalities at the input layer, building LLM trading agents that place news, prices, and charts into a language model's context and rely on its reasoning for decisions; (2) Specialized multimodal architectures develop dedicated modules to combine modality representations to exploit complementary information—information one modality carries but another lacks; (3) Large-scale pre-training builds temporal foundation models to capture rich temporal patterns from prices and technical signals.

Despite differing approaches, these lines mainly emphasize aggregating information to exploit complementary evidence rather than explicitly modeling whether modality contributions reinforce or cancel a shared predictive direction. The authors decompose multimodal information into three components: (1) Corroboration : task-conditioned, non-cancelling support from heterogeneous modalities for a common predictive direction—cross-view corroboration is less susceptible to modality-specific noise; (2) Complementarity : information contributed by one modality beyond others; (3) Conflict : modalities providing opposing support for the predictive direction. Corroboration is an overlooked but critical factor in financial trading—exploitable signals are weak, each view has high inherent noise, single-view predictions are fragile, while cross-view agreement provides a more stable trading basis.

Problem Formulation

Existing methods entangle shared structure—whether supportive evidence or cross-modal conflict—into a single fused vector, making latent corroboration implicit and hard to recover. Two core questions arise: (1) How to explicitly model multimodal corroboration instead of entangling it in a fused vector? (2) How to maintain reliable corroboration when individual modalities become noisy or unavailable?

Formally, given four time-aligned financial modalities (market, technical, news, sentiment) over a T-day historical window, the model predicts the next-day closing price direction of an asset: y_t = I[p_{t+1}^c ≥ p_t^c], where p_t^c is the adjusted closing price on day t, y_t = 1 for a bullish signal and y_t = 0 for a bearish signal.

Method

CoLAS comprises four modules: Multimodal Embedding , Corroboration Mining , Robust Prediction Layer , and Backtesting . Its core premise: when heterogeneous financial modalities corroborate a trading signal, that signal is reliable; when it relies on only a single modality, it is fragile.

4.1 Multimodal Encoders and Shared Alignment Space

Financial modalities have different statistical structures, so modality-aware networks are used instead of a single shared encoder: Market/Technical encoders are LSTMs summarizing temporal dynamics of both streams; News/Sentiment encoders aggregate pre-computed daily embeddings within the window plus a lightweight projection head. Each modality m's embedding is mapped via a modality-specific alignment mapping ψ_m to a shared alignment space ℝ^{d_c} and ℓ_2-normalized to the unit hypersphere, then column-stacked into a per-instance matrix V = [v̅_{m1}, v̅_{m2}, ..., v̅_{mk}] ∈ ℝ^{d_c × k}, k = |M|. Each column of V lies on the unit hypersphere, giving modality representations a unified scale and direct comparability, laying the foundation for extracting shared spectral structure and signed modality support.

4.2 Corroboration Mining

Multimodal corroboration is represented by two factors: spectral concentration and net signed support .

Spectral Concentration Enhancement: Singular Value Maximization. For each instance matrix V_i, SVD is performed; singular values are treated as logits and a softmax objective concentrates spectral energy into the dominant component: L_SVM = -(1/B) Σ log [ exp(λ_{i,1}/γ) / Σ exp(λ_{i,j}/γ) ]. Because each column of V_i is unit-normalized, total spectral energy is fixed at Σ λ_{i,j}^2 = k. Under this constraint, shifting energy from non-dominant to dominant components reduces the loss. The spectral concentration factor is defined as ρ_i = λ_{i,1}^2 / k.

Signed Corroboration Signal. Modality m's signed contribution along the dominant direction is a_{i,m} = p_{i,1}^⊤ v̅_{i,m}, where positive values indicate support and negative values indicate opposition. The corroboration signal is the projection of the modality mean onto the dominant direction z_i, whose strength factorizes as ‖z_i‖_2^2 = ρ_i · δ_i, where ρ_i measures spectral concentration and δ_i ∈ [0,1] measures net signed support. Therefore, a strong corroboration signal requires both a dominant shared component and limited cancellation among modality contributions —this is the fundamental difference between CoLAS and simple fusion methods.

Per-Instance Regularization. Using the SVM objective alone would collapse all instances to the same direction. A contrastive term regularizes corroboration signals across instances, suppressing similarity between different instances' corroboration signals to preserve cross-instance discriminability.

4.3 Robust Prediction Layer

Financial modality reliability varies significantly across trading days. CoLAS enhances robustness by training the model to produce consistent outputs on clean and perturbed samples (one modality corrupted). Each mini-batch randomly selects one modality and corrupts it via masking (zeroing) or additive Gaussian noise: X̃_a = Π(X_a) = { 0 (mask) or X_a + ε, ε ~ N(0, σ_ε^2 I) }. The robustness loss penalizes the difference (MSE) between clean and perturbed views at both the decision layer and the representation layer; the first term stabilizes prediction logits, the second preserves the corroboration signal under modality degradation.

4.4 Joint Optimization Objective

L_CoLAS = L_cls + β_1·L_SVM + β_2·L_IR + β_3·L_rob, where L_cls is binary cross-entropy classification loss, and β_1, β_2, β_3 control each objective's contribution. Since k ≪ d_c, recovering the dominant singular triplet from the compact Gram matrix yields a total mini-batch overhead of O(B·d_c·k^2 + B^2·d_c), dominated by the modality encoders.

Experiments

Experimental Setup

Datasets: Six asset-specific datasets—five U.S. stocks (AAPL, AMZN, GOOG, MSFT, TSLA) and Bitcoin (BTCUSD). Each dataset contains four time-aligned modalities: Market (Yahoo Finance: OHLCV), Technical (Finnhub: RSI, volatility, P/E, etc.), News (Alpaca: corporate events, announcements, etc.), Sentiment (Alpha Vantage: investor reactions, analyst opinions, etc.). Time splits: Training 2023.10–2024.09, Validation 2024.10–2025.03, Test 2025.04–2025.09. Baselines (16): Rule-based strategies (B&H, MACD, ZMR, SMA), single-modality models (LSTM, Transformer, DQN, PPO, Kronos temporal foundation model), general multimodal LLMs (Qwen3-8B, DeepSeek-R1-0528, Llama4-Scout-17B), financial multimodal LLMs (FinAgent, TradingAgents, DeepFund, VTA). 8×A100 GPU, all experiments repeated 5 times with mean reported.

Main Results

CoLAS achieves the best ARR and SR on all six datasets. U.S. Stocks: AAPL ARR 67.79% (SR 1.47, strongest baseline SR only 0.93); AMZN 56.98% (SR 1.25); GOOG 124.68% (SR 2.46); MSFT 97.16% (SR 2.73, vs VTA 2.42); TSLA 159.86% (SR 1.89). Cryptocurrency: BTCUSD ARR 84.64% (SR 2.65, vs VTA 1.92), a 23.0% ARR improvement over the runner-up VTA. CoLAS improves annualized return by over 25% on average relative to the respective strongest baseline across the six main datasets, and achieves the highest Sharpe ratio on all six. The simultaneous improvement in ARR and SR indicates the return gains do not come at the cost of worse risk-adjusted performance—CoLAS improves the return-risk trade-off itself.

Ablation Study

Removing SVM (singular value maximization): Largest ARR drop (AAPL 67.79%→59.54%, BTCUSD 84.64%→74.88%); dominant component concentration primarily contributes profitability. Removing IR (per-instance regularization): Both ARR and SR drop (AAPL ARR 60.49%, SR 1.30); loss of cross-instance discriminability. Removing RPL (robust prediction layer): Largest SR drop (AAPL 1.47→1.29, BTCUSD 2.65→2.14) but smallest ARR drop, indicating RPL mainly contributes risk-adjusted performance and robustness under modality degradation.

Corroboration Enhancement Analysis

Define "majority-opposing instances" as those where more than half of modality pairs have negative cosine similarity. Before joint optimization, 69%–96% of instances exhibit majority pairwise opposition; after joint optimization, this proportion falls to no more than 0.8% . This shows optimization drastically reduces inter-modality representational opposition, thereby enhancing the corroboration signal.

Extended Metrics and Hyperparameter Analysis

On four extended metrics (CR, MDD, SoR, CalR), CoLAS performs strongly. On BTCUSD, maximum drawdown is only 6.27%, Calmar ratio 13.50 (vs VTA 7.91). Hyperparameter analysis shows: β_1 ∈ {0.2, 0.5} and β_3 ∈ {0.5, 1.0} work best, γ = 0.10 is optimal. Regarding window size, CoLAS remains stable across three window sizes (AAPL ARR 65.49%–67.79%), while baselines like Kronos are highly sensitive to window size (ARR 12.79%–38.38%).

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

deep learningmultimodal learningstock predictionspectral analysisfinancial tradingcryptocurrency tradingcorroboration miningrobust prediction
Bighead's Algorithm Notes
Written by

Bighead's Algorithm Notes

Focused on AI applications in the fintech sector

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.