Machine Learning Algorithms & Natural Language Processing
Aug 10, 2026 · Artificial Intelligence
Can Models Keep Getting Stronger After Deployment? SERPO Enables Label‑Free Self‑Evolution on Open‑Ended Tasks
The SERPO framework replaces answer‑voting with a self‑evolving rubric, allowing test‑time reinforcement learning to improve large language models on open‑ended tasks without any human labels, and demonstrates up to 73% of privileged‑supervision gains on benchmarks such as HealthBench and ResearchQA.
AI EvaluationHealthBenchOpen-Ended Tasks
0 likes · 15 min read
