Bighead's Algorithm Notes
Sep 15, 2026 · Artificial Intelligence
FinSMART: Market-Aligned RL Trains LLMs on Real Returns, Beating FinDPO by 220%
FinSMART introduces a market-aligned reinforcement learning framework that trains LLMs for financial sentiment analysis using realized market returns instead of human annotations, employing GRPO with a dual-filter trading reward to achieve 220% higher cumulative returns than the previous state-of-the-art FinDPO.
FinSMARTGRPOLLM
0 likes · 23 min read
