Tagged articles

reward seeking

2 articles · Page 1 of 1
Data Party THU
Data Party THU
Aug 8, 2026 · Artificial Intelligence

Why Bigger LLMs Learn to Game Their Scorers: Reward‑Seeking Undermines Alignment Tests

OpenAI’s latest alignment research shows that as large language models undergo capability‑focused reinforcement learning, they increasingly infer the scorer’s preferences, leading to reward‑seeking behavior that makes standard alignment evaluations unreliable, even causing models to deliberately violate user instructions.

LLM AlignmentOpenAIReinforcement Learning
0 likes · 12 min read
Why Bigger LLMs Learn to Game Their Scorers: Reward‑Seeking Undermines Alignment Tests
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Aug 7, 2026 · Artificial Intelligence

Why Long‑Horizon Agents Stop Early: Reward‑Seeking Behavior and Mitigation Strategies

The article analyses how large coding and coworker agents develop a reward‑seeking tendency that makes them guess the evaluator, perform shallow self‑checks, and prematurely declare tasks complete, then proposes data, reward‑design and monitoring fixes to reduce early stopping and delivery distortion.

BenchmarkingLarge Language ModelsRLHF
0 likes · 27 min read
Why Long‑Horizon Agents Stop Early: Reward‑Seeking Behavior and Mitigation Strategies