Tagged articles

ByteDance Seed

3 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 18, 2026 · Artificial Intelligence

Self-Developing Agents: Three Benchmarks Reveal Why AI Struggles to Self-Improve

ByteDance Seed and TokenWave introduce three benchmarks—ASPIRE, S³Gym, and HarnessDev—to evaluate whether AI agents can autonomously form goals, learn from experience, and retain improvements, showing that current agents struggle to translate self-assessment into lasting capability gains.

AI agentsASPIREByteDance Seed
0 likes · 12 min read
Self-Developing Agents: Three Benchmarks Reveal Why AI Struggles to Self-Improve
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 16, 2026 · Artificial Intelligence

AI Solves Tests But Can't Self-Evolve: ByteDance Seed's Three RSI Benchmarks

ByteDance Seed and TokenWave introduce ASPIRE, S³Gym, and HarnessDev — three benchmarks that test whether AI agents can autonomously select learning goals, distill experience into improved decisions, and persistently upgrade their own execution systems without human-provided verification.

AI agentsASPIREBenchmarks
0 likes · 14 min read
AI Solves Tests But Can't Self-Evolve: ByteDance Seed's Three RSI Benchmarks
Machine Heart
Machine Heart
Sep 16, 2026 · Artificial Intelligence

Why Agents Struggle to Self-Evolve: Three Benchmarks for True Recursive Improvement

ByteDance Seed and collaborators introduce ASPIRE, S³Gym, and HarnessDev benchmarks to study how agents learn from vague goals, self-evaluate actions, and persist improvements, revealing that current agents overfit to proxy feedback and fail to convert self-judgment into lasting capability gains.

AI agentsASPIREAgent Benchmarks
0 likes · 11 min read
Why Agents Struggle to Self-Evolve: Three Benchmarks for True Recursive Improvement