Why VLMs Miss Reference Images and How RefCaptioner Aligns Them Precisely
RefCaptioner tackles the blind spot of existing video-language models by jointly grounding video captions to multiple reference images, using a dual‑reward fine‑tuning scheme and a new benchmark (MRVBench) that evaluates factual accuracy, image selection, and grounding robustness across up to 22 reference images per video.
