OpenRSI Index: Testing AI Research Agents on Real Tasks with 10K+ GPU Hours

OpenRSI Index evaluates AI agents on real open-source research projects — improving Marin's optimizer, Qwen's RL tasks, and GPIC image generation — using thousands of H100 hours, with isolated workspaces and independent verification to ensure reproducible progress beyond benchmark gaming.

Machine Heart
Machine Heart
Machine Heart
OpenRSI Index: Testing AI Research Agents on Real Tasks with 10K+ GPU Hours

Background: Benchmark Saturation vs. Real Research

As large language models approach saturation on standard benchmarks, a critical question emerges: does leaderboard progress translate to genuine research advancement? The article argues that "benchmark gaming" is not equivalent to effective research, and completing a single task does not demonstrate the ability to sustain research progress. Model evolution — from large-scale pre-training to RLHF — has always relied on valid feedback signals. When research moves toward recursive self-improvement (RSI), a key bottleneck appears: who defines the real, valid environmental feedback when a model tries to improve another model or itself?

OpenRSI Index: Real-World Research Tasks at Scale

OpenRSI Index addresses this by sourcing tasks from influential open-source research projects (Marin, Qwen, GPIC) and letting agents improve upon existing models, code, data, and training recipes. These tasks require actual model training experiments, consuming thousands to tens of thousands of H100·hours. The platform's website (https://index.openrsi.foundation/index.html) and GitHub repository (https://github.com/OpenRSI-Foundation/OpenRSI-Index) provide open access to tasks, evaluation framework, verification tools, and run trajectories.

Three Representative Tasks

Pre-training: Improving Marin's optimizer. Marin-Scaling-Ladder tasks the agent with applying the same optimizer rules across six model scales (550M to 2.545B parameters) to test whether a method remains effective as model size changes. Single run: 18,432 H100·hours .

Post-training: Designing better RL tasks for Qwen. Qwen-122B-RL-Merge asks the agent to generate verifiable training tasks that improve Qwen's performance on math, code, and other domains. Single run: 20,864 H100·hours .

Visual generation: Learning better generation from GPIC data. Starting from a shared model checkpoint, the agent explores architecture, training objective, or recipe improvements; each candidate retrains on the same 10M images at most once. Single run: 5,815 H100·hours .

Together, these tasks probe three questions: Can a method scale across model sizes? Can an agent organize experiments based on feedback? Can a single improvement be clearly attributed?

RSI-Harness: Dual-Loop Verification Mechanism

To counter the "suspicion chain" in research evaluation (overfitting to test distribution, chance amplification of noise, irreproducibility), OpenRSI Index enforces a strict separation between the agent's experimentation workspace and the final acceptance environment. The mechanism operates in two loops:

Inner research loop: The agent persistently modifies code and runs experiments in a workspace, iterating based on feedback.

Outer verification loop: Each formal submission freezes a snapshot, which is then re-evaluated in a freshly launched, independent Judge environment. Private tests are never exposed to the agent; file changes in the evaluation environment do not return to the workspace, but scores and prescribed feedback do. Hard violations (reward tampering, test leakage, budget overrun) receive zero score.

The article includes a diagram of this workflow (

RSI-Harness workspace, independent acceptance, and feedback flow
RSI-Harness workspace, independent acceptance, and feedback flow

).

Independent Verification and Reproducibility

The original authors' solutions are re-run in the same task environment as baselines. This design does not demand a "one-shot correct answer"; agents may trial-and-error. However, any claimed improvement must hold not only in the agent's own logs but also in an environment it cannot control. The article notes that verifiability and traceability give researchers a concrete basis to assess whether cited work is truly trustworthy.

Public experiment logs for the GPIC text-to-image task show the agent repeatedly adjusting training recipes and improving based on FD-DINOv2 scores (lower is better) over time (

GPIC task public experiment measurements: research time vs. FD-DINOv2 score
GPIC task public experiment measurements: research time vs. FD-DINOv2 score

).

Mentor Team and Community-Driven Expansion

OpenRSI Index is backed by a distinguished mentor board including Luke Zettlemoyer (UW, co-author of RoBERTa, BART, Toolformer), Karthik Narasimhan (Princeton, co-author of GPT-1, ReAct, SWE-bench, SWE-agent), Yejin Choi (Stanford, HellaSwag, ATOMIC, COMET), and Dawn Song (UC Berkeley, MMLU, MATH, ALE, CyberGym, ExploitGym), plus advisors spanning language, multimodal, systems, and security research (

OpenRSI Index mentor lineup
OpenRSI Index mentor lineup

).

RSI-Anything: Lowering Barriers for Task Contribution

The team believes research is a positive-sum game. Through RSI-Anything, researchers can convert their open research into runnable RSI tasks by conversing with an AI for about one hour to confirm key settings; the AI then auto-generates the proposal and builds the environment. Accepted tasks (after review and reproduction) earn contributors co-authorship. The platform also seeks compute partners and domain leads in foundation models, biomedicine, law, 3D vision, scientific computing, etc. (

RSI-Anything task contribution pipeline
RSI-Anything task contribution pipeline

).

The overarching goal: From the community, for the community — letting more researchers bring real problems in, making successes, failures, and intermediate processes inspectable, reusable, and extensible, so that RSI belongs to everyone and everyone participates in building RSI evaluation standards.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

RLHFopen-source researchbenchmark evaluationGPU computerecursive self-improvementindependent verificationAI research agentsOpenRSI Index
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.