VLX-VR Breaks Long-Video 'Duration Curse', Tops MINERVA Benchmark
Om AI's VLX-VR model achieves 78.8% accuracy on Google DeepMind's MINERVA long-video reasoning benchmark, outperforming Gemini 3.5 and GPT-4.1, and uniquely improves accuracy on videos over 15 minutes, reversing the typical performance drop with longer videos.
VLX-VR: End-to-End Hour-Long Video Reasoning
Om AI's latest model, VLX-VR , targets complex reasoning over dynamic long videos. It is the newest member of the VLX on-device streaming multimodal series , which already includes VLX-Flow (continuous perception), VLX-Seek (precise localization), and VLX-Go (action decision). VLX-VR extends visual reasoning to dynamic long videos , supporting cross-temporal complex reasoning without manual video clipping.
MINERVA Benchmark: The Long-Video Reasoning Challenge
Google DeepMind's MINERVA benchmark comprises 1000+ questions across 200+ videos ranging from under 2 minutes to over 1.5 hours, spanning sports, short films, tutorials, travel, and life skills. Unlike traditional VideoQA focused on content recognition, MINERVA questions require integrating information from multiple temporal segments — e.g., an action occurring tens of minutes earlier becomes critical evidence for a later question.
Benchmark Results: Reversing the "Duration Curse"
On MINERVA, VLX-VR achieves 78.8% overall video QA accuracy , surpassing Google Gemini 3.5, OpenAI GPT-4.1, and other flagship models, setting a new public evaluation record. Human accuracy stands at 92.54%; the previous best reported model (Gemini 2.5 Pro Thinking) scored 66.20%.
Crucially, VLX-VR improves with video length :
Under 5 minutes: VLX-VR 76.70% vs. Gemini 2.5 Pro Thinking 68.87% (+7.8 pp)
15+ minutes: VLX-VR 80.92% vs. Gemini 2.5 Pro Thinking 57.97% and GPT-4.1 47.25% (lead expands to ~23 pp)
Across three duration buckets, VLX-VR's cross-duration accuracy variance (CDAV) is only 2.97 (percentage points²) , corresponding to a standard deviation of ~1.72 pp, demonstrating stable performance regardless of video length.
VLX-VR vs. mainstream models on MINERVA accuracy
Accuracy across video durations
VLX-VR's performance trend vs. other models
Reasoning Trace Consistency: Verifiable Process
MINERVA provides human-annotated reasoning traces (avg. 92 words, 99.6% contain timestamps, avg. 4 time points per trace). On correctly answered samples, VLX-VR's generated reasoning traces match human references at 96.2% consistency . This traceability allows developers to pinpoint whether errors stem from temporal localization, visual understanding, or logical reasoning — critical for robotics, smart glasses, and video analytics.
Capability Breakdown
Reading Comprehension: 90.76% (extracts subtitles, signs, on-screen text)
Situational Awareness: 90.32% (understands scene, behavior, environment state)
Temporal Reasoning: 85.95% (orders events, links cross-temporal segments)
Numerical Reasoning: 85.71% (handles quantities and calculations in video)
Capability-dimension accuracy on MINERVA
Real-World Open-Video Tests
Anomaly detection in basketball arena: Given a basketball court video where players perform soccer kicks and headers, VLX-VR identified 5 anomalous actions with timestamps , while Gemini-3.1-pro misjudged the video as reversed and built analysis on that false premise.
Anomaly detection comparison
Chess game state tracking: VLX-VR located the first check, then continuously tracked subsequent moves to determine how many moves White made before the endgame — demonstrating persistent state maintenance across time.
Chess game cross-temporal reasoning
From VLM-R1 to VLX-VR: Extending Reasoning to Time
In February 2025, Om AI open-sourced VLM-R1 , introducing R1-style reinforcement learning reasoning to vision. VLX-VR extends that visual reasoning capability to dynamic long video , adding the critical dimension of time . The VLX series now covers:
VLX-Flow: Continuous perception & real-time response
VLX-Seek: Precise target localization & spatial relations
VLX-Go: Connecting visual understanding to action
VLX-VR: Watch full videos, perform deep reasoning over longer time scales
VLX-VR is now available on the OmAgent official experience platform for public testing.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Machine Learning Algorithms & Natural Language Processing
Focused on frontier AI technologies, empowering AI researchers' progress.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
