Machine Heart
Aug 11, 2026 · Artificial Intelligence
Why VLMs Miss Reference Images and How RefCaptioner Aligns Them Precisely
RefCaptioner tackles the blind spot of existing video-language models by jointly grounding video captions to multiple reference images, using a dual‑reward fine‑tuning scheme and a new benchmark (MRVBench) that evaluates factual accuracy, image selection, and grounding robustness across up to 22 reference images per video.
MRVBenchQwen3-VL-8BRefCaptioner
0 likes · 13 min read
