Vision-OPD Enables Multimodal LLMs to See Fine Details in a Single Forward Pass
Vision-OPD introduces an online self‑distillation framework that lets a 9B multimodal LLM internalize fine‑grained visual evidence from whole images, achieving state‑of‑the‑art results on six detailed visual‑understanding benchmarks without extra inference tools.
