AgentHOI: Training‑Free Human‑Object Interaction Detection with Multimodal LLMs
The paper introduces AgentHOI, a framework that eliminates the need for HOI‑specific annotations by decomposing detection into context‑aware semantic reasoning and multifaceted spatial localization, leveraging GPT‑4o and GroundingDINO, and demonstrates strong zero‑shot and robustness performance on HICO‑DET and SWIG‑HOI benchmarks.
