Tagged articles

MRVBench

1 articles · Page 1 of 1
Machine Heart
Machine Heart
Aug 11, 2026 · Artificial Intelligence

Why VLMs Miss Reference Images and How RefCaptioner Aligns Them Precisely

RefCaptioner tackles the blind spot of existing video-language models by jointly grounding video captions to multiple reference images, using a dual‑reward fine‑tuning scheme and a new benchmark (MRVBench) that evaluates factual accuracy, image selection, and grounding robustness across up to 22 reference images per video.

MRVBenchQwen3-VL-8BRefCaptioner
0 likes · 13 min read
Why VLMs Miss Reference Images and How RefCaptioner Aligns Them Precisely