Tagged articles

multimodal VLM

4 articles · Page 1 of 1
Machine Heart
Machine Heart
Aug 11, 2026 · Artificial Intelligence

Why VLMs Miss Reference Images and How RefCaptioner Aligns Them Precisely

RefCaptioner tackles the blind spot of existing video-language models by jointly grounding video captions to multiple reference images, using a dual‑reward fine‑tuning scheme and a new benchmark (MRVBench) that evaluates factual accuracy, image selection, and grounding robustness across up to 22 reference images per video.

MRVBenchQwen3-VL-8BRefCaptioner
0 likes · 13 min read
Why VLMs Miss Reference Images and How RefCaptioner Aligns Them Precisely
vivo Internet Technology
vivo Internet Technology
Jul 8, 2026 · Artificial Intelligence

VeraRetouch: A Multi‑Task Inference Framework for Photo Editing (SIGGRAPH 2026)

VeraRetouch is a lightweight, fully differentiable photo‑retouching framework that runs on mobile devices, uses a 0.6B visual‑language model as a “retouch brain”, introduces a differentiable Retouch Renderer with three control dimensions, leverages a million‑scale AetherRetouch‑1M+ dataset, and achieves state‑of‑the‑art quality, speed, and deployability across automatic, style‑guided, and parameter‑driven editing tasks.

AIdifferentiable renderingimage retouching
0 likes · 13 min read
VeraRetouch: A Multi‑Task Inference Framework for Photo Editing (SIGGRAPH 2026)
DataFunSummit
DataFunSummit
Jul 7, 2026 · Artificial Intelligence

Ant Group OpAgent: Online RL‑Powered Open‑Domain Browser Automation Agent

The article details Ant Group's OpAgent, an open‑domain browser automation agent that overcomes perception, timeliness, and implicit interaction challenges through a three‑stage pipeline of multi‑task supervised fine‑tuning, online reinforcement learning, and a four‑module Planner‑Grounder‑Reflector‑Summarizer architecture, achieving a 71.6% Pass@1 score on WebArena and releasing all code and models publicly.

Agent ArchitectureOpAgentWebArena
0 likes · 15 min read
Ant Group OpAgent: Online RL‑Powered Open‑Domain Browser Automation Agent
Machine Heart
Machine Heart
May 6, 2026 · Artificial Intelligence

PromptEcho: Leveraging Frozen Multimodal Models for High‑Quality Text‑to‑Image Rewards Without Labels

PromptEcho computes a continuous reward for text‑to‑image generation by measuring how well a frozen vision‑language model can reconstruct the original prompt from the generated image, eliminating the need for annotated data or a trained reward model and outperforming prior methods across multiple benchmarks.

PromptEchobenchmarkdense captioning
0 likes · 10 min read
PromptEcho: Leveraging Frozen Multimodal Models for High‑Quality Text‑to‑Image Rewards Without Labels