OpenRSI Index: Testing AI Research Agents on Real Tasks with 10K+ GPU Hours
OpenRSI Index evaluates AI agents on real open-source research projects — improving Marin's optimizer, Qwen's RL tasks, and GPIC image generation — using thousands of H100 hours, with isolated workspaces and independent verification to ensure reproducible progress beyond benchmark gaming.
Background: Benchmark Saturation vs. Real Research
As large language models approach saturation on standard benchmarks, a critical question emerges: does leaderboard progress translate to genuine research advancement? The article argues that "benchmark gaming" is not equivalent to effective research, and completing a single task does not demonstrate the ability to sustain research progress. Model evolution — from large-scale pre-training to RLHF — has always relied on valid feedback signals. When research moves toward recursive self-improvement (RSI), a key bottleneck appears: who defines the real, valid environmental feedback when a model tries to improve another model or itself?
OpenRSI Index: Real-World Research Tasks at Scale
OpenRSI Index addresses this by sourcing tasks from influential open-source research projects (Marin, Qwen, GPIC) and letting agents improve upon existing models, code, data, and training recipes. These tasks require actual model training experiments, consuming thousands to tens of thousands of H100·hours. The platform's website (https://index.openrsi.foundation/index.html) and GitHub repository (https://github.com/OpenRSI-Foundation/OpenRSI-Index) provide open access to tasks, evaluation framework, verification tools, and run trajectories.
Three Representative Tasks
Pre-training: Improving Marin's optimizer. Marin-Scaling-Ladder tasks the agent with applying the same optimizer rules across six model scales (550M to 2.545B parameters) to test whether a method remains effective as model size changes. Single run: 18,432 H100·hours .
Post-training: Designing better RL tasks for Qwen. Qwen-122B-RL-Merge asks the agent to generate verifiable training tasks that improve Qwen's performance on math, code, and other domains. Single run: 20,864 H100·hours .
Visual generation: Learning better generation from GPIC data. Starting from a shared model checkpoint, the agent explores architecture, training objective, or recipe improvements; each candidate retrains on the same 10M images at most once. Single run: 5,815 H100·hours .
Together, these tasks probe three questions: Can a method scale across model sizes? Can an agent organize experiments based on feedback? Can a single improvement be clearly attributed?
RSI-Harness: Dual-Loop Verification Mechanism
To counter the "suspicion chain" in research evaluation (overfitting to test distribution, chance amplification of noise, irreproducibility), OpenRSI Index enforces a strict separation between the agent's experimentation workspace and the final acceptance environment. The mechanism operates in two loops:
Inner research loop: The agent persistently modifies code and runs experiments in a workspace, iterating based on feedback.
Outer verification loop: Each formal submission freezes a snapshot, which is then re-evaluated in a freshly launched, independent Judge environment. Private tests are never exposed to the agent; file changes in the evaluation environment do not return to the workspace, but scores and prescribed feedback do. Hard violations (reward tampering, test leakage, budget overrun) receive zero score.
The article includes a diagram of this workflow (
).
Independent Verification and Reproducibility
The original authors' solutions are re-run in the same task environment as baselines. This design does not demand a "one-shot correct answer"; agents may trial-and-error. However, any claimed improvement must hold not only in the agent's own logs but also in an environment it cannot control. The article notes that verifiability and traceability give researchers a concrete basis to assess whether cited work is truly trustworthy.
Public experiment logs for the GPIC text-to-image task show the agent repeatedly adjusting training recipes and improving based on FD-DINOv2 scores (lower is better) over time (
).
Mentor Team and Community-Driven Expansion
OpenRSI Index is backed by a distinguished mentor board including Luke Zettlemoyer (UW, co-author of RoBERTa, BART, Toolformer), Karthik Narasimhan (Princeton, co-author of GPT-1, ReAct, SWE-bench, SWE-agent), Yejin Choi (Stanford, HellaSwag, ATOMIC, COMET), and Dawn Song (UC Berkeley, MMLU, MATH, ALE, CyberGym, ExploitGym), plus advisors spanning language, multimodal, systems, and security research (
).
RSI-Anything: Lowering Barriers for Task Contribution
The team believes research is a positive-sum game. Through RSI-Anything, researchers can convert their open research into runnable RSI tasks by conversing with an AI for about one hour to confirm key settings; the AI then auto-generates the proposal and builds the environment. Accepted tasks (after review and reproduction) earn contributors co-authorship. The platform also seeks compute partners and domain leads in foundation models, biomedicine, law, 3D vision, scientific computing, etc. (
).
The overarching goal: From the community, for the community — letting more researchers bring real problems in, making successes, failures, and intermediate processes inspectable, reusable, and extensible, so that RSI belongs to everyone and everyone participates in building RSI evaluation standards.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
