VLX-VR Breaks Long-Video 'Duration Curse', Tops MINERVA Benchmark

Om AI's VLX-VR model achieves 78.8% accuracy on Google DeepMind's MINERVA long-video reasoning benchmark, outperforming Gemini 3.5 and GPT-4.1, and uniquely improves accuracy on videos over 15 minutes, reversing the typical performance drop with longer videos.

Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
VLX-VR Breaks Long-Video 'Duration Curse', Tops MINERVA Benchmark

VLX-VR: End-to-End Hour-Long Video Reasoning

Om AI's latest model, VLX-VR , targets complex reasoning over dynamic long videos. It is the newest member of the VLX on-device streaming multimodal series , which already includes VLX-Flow (continuous perception), VLX-Seek (precise localization), and VLX-Go (action decision). VLX-VR extends visual reasoning to dynamic long videos , supporting cross-temporal complex reasoning without manual video clipping.

MINERVA Benchmark: The Long-Video Reasoning Challenge

Google DeepMind's MINERVA benchmark comprises 1000+ questions across 200+ videos ranging from under 2 minutes to over 1.5 hours, spanning sports, short films, tutorials, travel, and life skills. Unlike traditional VideoQA focused on content recognition, MINERVA questions require integrating information from multiple temporal segments — e.g., an action occurring tens of minutes earlier becomes critical evidence for a later question.

Benchmark Results: Reversing the "Duration Curse"

On MINERVA, VLX-VR achieves 78.8% overall video QA accuracy , surpassing Google Gemini 3.5, OpenAI GPT-4.1, and other flagship models, setting a new public evaluation record. Human accuracy stands at 92.54%; the previous best reported model (Gemini 2.5 Pro Thinking) scored 66.20%.

Crucially, VLX-VR improves with video length :

Under 5 minutes: VLX-VR 76.70% vs. Gemini 2.5 Pro Thinking 68.87% (+7.8 pp)

15+ minutes: VLX-VR 80.92% vs. Gemini 2.5 Pro Thinking 57.97% and GPT-4.1 47.25% (lead expands to ~23 pp)

Across three duration buckets, VLX-VR's cross-duration accuracy variance (CDAV) is only 2.97 (percentage points²) , corresponding to a standard deviation of ~1.72 pp, demonstrating stable performance regardless of video length.

VLX-VR vs. mainstream models on MINERVA accuracy comparison
VLX-VR vs. mainstream models on MINERVA accuracy comparison

VLX-VR vs. mainstream models on MINERVA accuracy

Multiple models' MINERVA accuracy across video durations
Multiple models' MINERVA accuracy across video durations

Accuracy across video durations

VLX-VR shows opposite performance trend as video length increases
VLX-VR shows opposite performance trend as video length increases

VLX-VR's performance trend vs. other models

Reasoning Trace Consistency: Verifiable Process

MINERVA provides human-annotated reasoning traces (avg. 92 words, 99.6% contain timestamps, avg. 4 time points per trace). On correctly answered samples, VLX-VR's generated reasoning traces match human references at 96.2% consistency . This traceability allows developers to pinpoint whether errors stem from temporal localization, visual understanding, or logical reasoning — critical for robotics, smart glasses, and video analytics.

Capability Breakdown

Reading Comprehension: 90.76% (extracts subtitles, signs, on-screen text)

Situational Awareness: 90.32% (understands scene, behavior, environment state)

Temporal Reasoning: 85.95% (orders events, links cross-temporal segments)

Numerical Reasoning: 85.71% (handles quantities and calculations in video)

VLX-VR accuracy across MINERVA capability dimensions
VLX-VR accuracy across MINERVA capability dimensions

Capability-dimension accuracy on MINERVA

Real-World Open-Video Tests

Anomaly detection in basketball arena: Given a basketball court video where players perform soccer kicks and headers, VLX-VR identified 5 anomalous actions with timestamps , while Gemini-3.1-pro misjudged the video as reversed and built analysis on that false premise.

VLX-VR vs. Gemini-3.1-pro on basketball anomaly detection
VLX-VR vs. Gemini-3.1-pro on basketball anomaly detection

Anomaly detection comparison

Chess game state tracking: VLX-VR located the first check, then continuously tracked subsequent moves to determine how many moves White made before the endgame — demonstrating persistent state maintenance across time.

VLX-VR cross-temporal state reasoning in chess video
VLX-VR cross-temporal state reasoning in chess video

Chess game cross-temporal reasoning

From VLM-R1 to VLX-VR: Extending Reasoning to Time

In February 2025, Om AI open-sourced VLM-R1 , introducing R1-style reinforcement learning reasoning to vision. VLX-VR extends that visual reasoning capability to dynamic long video , adding the critical dimension of time . The VLX series now covers:

VLX-Flow: Continuous perception & real-time response

VLX-Seek: Precise target localization & spatial relations

VLX-Go: Connecting visual understanding to action

VLX-VR: Watch full videos, perform deep reasoning over longer time scales

VLX-VR is now available on the OmAgent official experience platform for public testing.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

video understandingmultimodal modelstemporal reasoningbenchmark resultslong video reasoningOm AIMINERVA benchmarkVLX-VR
Machine Learning Algorithms & Natural Language Processing
Written by

Machine Learning Algorithms & Natural Language Processing

Focused on frontier AI technologies, empowering AI researchers' progress.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.