FinSMART: Market-Aligned RL Trains LLMs on Real Returns, Beating FinDPO by 220%

FinSMART introduces a market-aligned reinforcement learning framework that trains LLMs for financial sentiment analysis using realized market returns instead of human annotations, employing GRPO with a dual-filter trading reward to achieve 220% higher cumulative returns than the previous state-of-the-art FinDPO.

Bighead's Algorithm Notes
Bighead's Algorithm Notes
Bighead's Algorithm Notes
FinSMART: Market-Aligned RL Trains LLMs on Real Returns, Beating FinDPO by 220%

Abstract

Financial sentiment analysis extracts market-relevant signals from unstructured text (news, earnings reports) and is a core component of algorithmic and event-driven trading strategies. Recent financial LLMs (FinBERT, FinGPT, FinLlama, FinDPO) have improved accuracy, but all remain trapped in a market-agnostic supervised learning paradigm — they rely on limited, static, human-annotated datasets and cannot adapt to evolving market conditions.

This paper proposes FinSMART — the first market-aligned reinforcement learning framework for financial sentiment analysis. Instead of human labels, it uses realized market returns as the training signal, optimizing the LLM's sentiment predictions via Group Relative Policy Optimization (GRPO). To handle the inherent noise, non-stationarity, and heavy-tailed distribution of financial returns, FinSMART designs a signal extraction pipeline that combines market-aware data filtering with a discrete asymmetric trading reward (dual-filter trading reward) to extract economically meaningful feedback from noise.

Experiments show FinSMART comprehensively outperforms existing state-of-the-art methods on profitability, risk-adjusted performance, and sentiment signal quality, with cumulative trading returns improving 220% over the strongest baseline . The framework also natively supports market-aware retraining: simply replace human annotations with newly observed articles and their realized market returns, and the model continuously adapts to market dynamics.

Background

Approximately 80% of financial market data is unstructured text. Effectively using LLMs to extract actionable insights from complex narratives and convert them into autonomous trading signals can yield significant information advantages. However, the diversity, subtlety, and domain specificity of financial text pose major challenges for reliable sentiment extraction.

Evolution of Financial Sentiment LLMs

FinBERT (2019) : Fine-tuned BERT on ~4,000 annotated samples. Improved general-model performance in finance but limited by model size (110M parameters), insensitive to numerals, and lacked robustness on complex sentences.

FinGPT / Instruct-FinGPT (2023) : Shifted to larger general LLMs (Llama-7B, Llama-2-13B) with supervised instruction tuning for finance. Not specifically optimized for sentiment analysis; only produced categorical sentiment (positive/negative/neutral) without quantifying intensity.

FinLlama (2024) : Combined SFT with Llama-2-7B and added a classification head for continuous sentiment scores. Changed the model's core function from next-token generation to classification, limiting compatibility with generation-dependent post-training techniques. SFT also risked memorizing training samples, hurting generalization to unseen data.

FinDPO (2025) : Applied Direct Preference Optimization (DPO) with a logit-to-score conversion mechanism to turn discrete sentiment predictions into continuous scores. Current SOTA for financial sentiment analysis. However, FinDPO still depends on static annotated datasets and cannot adapt to changing market conditions.

All existing methods share a common limitation: the market-agnostic supervised learning paradigm — models train on fixed annotated data then deploy, unable to learn from subsequent market outcomes. FinSMART's core insight: can we "close the loop" and let the model learn directly from market returns?

Problem Statement

Existing financial sentiment analysis faces three dilemmas:

Annotation bottleneck : Human annotation is expensive and scarce; datasets are limited (typically thousands to tens of thousands of samples) and cannot cover the full dynamics of the market.

Market irrelevance : Annotated data is fixed at annotation time and cannot reflect subsequent actual market movements. Models learn "dictionary-style" sentiment mappings but do not know which sentiment signals truly have predictive value.

Inability to adapt : Market environments constantly evolve (e.g., COVID-19 crash, interest-rate cycle shifts); static models cannot self-adapt.

The core question: can we develop a financial sentiment analysis framework that does not rely on static annotated data, learns directly from real-time market dynamics, and aligns text sentiment with subsequent market behavior?

Method

4.1 Theoretical Foundation: GRPO Optimization

FinSMART adopts Group Relative Policy Optimization (GRPO) as its optimization framework. Unlike traditional PPO, GRPO does not require a separate value model (critic); instead, it samples a group of G responses from the same prompt and constructs the optimization objective using within-group relative rewards.

Given policy LLM π<sub>θ</sub>, for prompt x sample G responses, each receiving reward r<sub>i</sub> = R(x, y<sub>i</sub>). The relative advantage is computed via within-group normalization:

A<sub>i</sub> = (r<sub>i</sub> - μ(r)) / (σ(r) + δ)

The policy is optimized via a clipped surrogate objective with KL-divergence regularization to prevent the policy from drifting too far from the reference model:

L<sub>GRPO</sub>(θ) = - E[ (1/G) Σ min{ ρ<sub>i</sub> · A<sub>i</sub>, clip(ρ<sub>i</sub>, 1-ε, 1+ε) · A<sub>i</sub> } - β · D<sub>KL</sub>(π<sub>θ</sub> ∥ π<sub>ref</sub>) ]

where ρ<sub>i</sub> = π<sub>θ</sub>(y<sub>i</sub>|x) / π<sub>θold</sub>(y<sub>i</sub>|x) is the policy ratio, ε controls the clipping range per update, and β weights the KL penalty.

Key advantage of GRPO : By learning from within-group relative ranking, it avoids dependence on absolute reward calibration, making it especially suitable for financial market scenarios where the reward signal is noisy and difficult to model explicitly.

4.2 Three-Stage Training Pipeline

FinSMART comprises three stages: Data Alignment & Signal Extraction , Dual-Filter Trading Reward , and GRPO Policy Optimization .

Stage 1: Data Alignment & Signal Extraction

Financial returns are extremely noisy, non-stationary, and heavy-tailed. FinSMART's core premise: RL learning from market feedback is effective only when (i) the article contains a decodable sentiment signal, and (ii) the article can be reliably aligned to realized market outcomes.

Data Collection : Financial news from The Motley Fool (TMF) and MarketWatch, Feb 2015 – Jun 2021, matched with daily S&P 500 constituent market data (Yahoo Finance), covering 1,672 trading days.

Named Entity Recognition (NER) : Financial articles often discuss multiple companies; naively linking to a single stock introduces reward misattribution. BERT-base-NER identifies primary organizational entities; only articles where the target company's confidence exceeds 98% are retained. This drastically reduces the risk of attributing market returns to irrelevant articles.

Sentiment Gating : The reference policy π<sub>ref</sub> (pre-trained LLM) acts as a sentiment gate — each article outputs a single label {Positive, Negative, Neutral} under a constrained classification prompt. Articles the model cannot label are discarded, preventing noise from entering the reward signal.

4.3 Dual-Filter Trading Reward

RL effectiveness hinges on reward signal quality. In financial markets, realized return is a challenging supervisory signal — naively rewarding the model by subsequent return exposes the policy to extremely low signal-to-noise ratio, causing unstable optimization and poor generalization. FinSMART designs a market-aligned reward function with two goals: (i) extract economically meaningful learning signals from noisy market outcomes, (ii) provide a stable optimization landscape to avoid mode collapse.

Given article x, the policy generates sentiment prediction d ∈ {-1, 0, +1} (negative/neutral/positive). Let α be the stock's idiosyncratic return on the publication day (realized stock return minus S&P 500 index return, proxied by SPY ETF), and r be the raw realized stock return. The ground-truth direction label is defined as:

y = +1, if α > τ and r > 0
y = -1, if α < -τ and r < 0
y = 0, otherwise

where τ = 0.5% is the minimum alpha threshold. The reward function is discrete and asymmetric:

R<sub>trade</sub>(d, y):
+2.0, if d = y and y ≠ 0   (correct direction)
+0.1, if d = y = 0           (correct neutral)
-1.5, if d = -y and y ≠ 0    (opposite direction)
-1.0, otherwise              (missed trade or false signal)

This design has two key properties:

Dual Filtering : Unlike standard classification rewards, a sentiment prediction must satisfy two independent market conditions to earn a positive reward. First, the predicted direction must match the realized return (ensuring the implied trade position is profitable). Second, the corresponding alpha return must exceed the minimum threshold τ (ensuring the price move is attributable to firm-specific information rather than broad market swings). This reduces the influence of spurious price movements, extracting more informative learning signals from inherently noisy data.

Asymmetry : The positive reward for correct predictions (+2.0) exceeds the penalty for incorrect predictions (-1.5), encouraging exploration during training and preventing the model from becoming overly conservative and collapsing to predominantly neutral predictions due to uncertain market feedback.

Training uses same-day returns; testing uses next-day returns : The separation between publication-day alpha and sentiment is ~5.0% on TMF (between positive/negative articles), dropping to only 0.3% the next day; MarketWatch falls from 2.3% to 0.3%. Pearson correlation drops from 0.41→0.03 (TMF) and 0.37→0.03 (MarketWatch). Therefore training uses same-day returns (richer signal), but all trading experiments use next-day returns to eliminate look-ahead bias.

4.4 GRPO Training Configuration

Base model: Llama-3-8B-Instruct . Sample G=8 completions for exploration, KL regularization β = 0.1. LoRA fine-tuning: rank r=16, scaling α=32, dropout 0.05. Trainable parameters only 13.6M (~0.17% of base model parameters). The entire RL pipeline runs on a single NVIDIA A6000 (48 GB) , training time 8 hours. Training set: articles before Dec 2018; out-of-sample test period: Jan 2019 – Jun 2021 (includes COVID-19 crash, testing robustness under high volatility).

4.5 Sentiment-Driven Portfolio Construction

FinSMART uses a causal LLM outputting discrete sentiment labels. Via the logit-to-score converter proposed in FinDPO, it leverages the model's internal output logits to quantify the strength of each sentiment prediction, obtaining continuous sentiment scores without architectural modification.

Portfolio construction process : Investable universe = S&P 500 constituents, dynamically adjusted daily (additions, delistings, M&A) to eliminate survivorship bias. Rank by sentiment score; top 35% long, bottom 35% short, equal-weight. Day-t article sentiment builds portfolio at t+1 open, held to t+2 open (next-day open-to-open return), eliminating look-ahead bias.

Experiments

Experimental Setup

Data Sources : The Motley Fool + MarketWatch, Feb 2015 – Jun 2021. Investable universe: S&P 500 constituents. ~325,000 out-of-sample articles.

Baselines (7 methods) : Dictionary-based (LMD, HIV-4, VADER), SFT fine-tuned LLMs (FinBERT, FinLlama), Preference-optimized LLM (FinDPO, current SOTA).

Evaluation Metrics : Cumulative return, Annualized return, Sharpe ratio, Sortino ratio, Calmar ratio, Rank IC. Risk-free rate R<sub>f</sub> = 0.

Main Results

FinSMART leads across all evaluation metrics.

vs. FinDPO (current SOTA, same base model Llama-3-8B)

Cumulative return: 264.9% vs 109.8% (+141%)

Annualized return: 91.5% vs 45.0% (more than double)

Sharpe ratio: 1.97 vs 1.12

Sortino ratio: 2.40 vs 1.44

Calmar ratio: 4.23 vs 1.48

Rank IC: 0.061 vs 0.053 (+15%)

vs. Other Baselines

Dictionary methods (HIV-4, VADER, LMD): all negative returns, negative Sharpe.

FinBERT: Cumulative 45.2%, Sharpe 0.67.

FinLlama: Cumulative 87.9%, Sharpe 0.98.

S&P 500 benchmark: Cumulative 69.3%, Sharpe 1.09.

Relative to strongest baseline, cumulative return improves 220% .

The Rank IC improvement is especially important — it measures the cross-sectional alignment between sentiment scores and next-day alpha returns, showing that FinSMART's sentiment signals themselves are more predictive , not just higher returns.

Periodic Market-Aligned Retraining

FinSMART's unique advantage is support for market-aware retraining. No human annotation needed; simply pair newly published articles with their realized market returns to generate new training data. An expanding-window retraining strategy is used: retrain every 6 months on cumulative data, for 4 iterations total.

Retraining vs. Static Model

Cumulative return: 406.2% vs 264.9%

Annualized return: 125.7% vs 91.5%

Sharpe: 2.41 vs 1.97

Sortino: 2.96 vs 2.40

Calmar: 5.65 vs 4.23

Rank IC: 0.065 vs 0.061

The retrained model consistently outperforms the static baseline throughout the evaluation period. An interesting finding: a moderate positive correlation (r = 0.72) exists between the number of new articles added each 6-month period and the retrained model's performance gain over the static baseline, indicating that richer input information amplifies the benefits of periodic retraining . This validates FinSMART's ability to continuously re-align the sentiment model to evolving market dynamics, producing incrementally stronger and more informative trading signals.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LLMLoRAreinforcement learningGRPOalgorithmic tradingfinancial sentiment analysisFinSMARTmarket-aligned training
Bighead's Algorithm Notes
Written by

Bighead's Algorithm Notes

Focused on AI applications in the fintech sector

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.