Six 360 AI Institute Papers Accepted at 2026 Top AI Conferences Signal a Shift Toward Precise, Controllable AI

In the first half of 2026, 360 AI Institute saw six papers accepted at ICLR, CVPR, ICML and ECCV, covering multimodal understanding, agents, image generation and editing, while Chinese contributions reached 43.7% of ICLR submissions, highlighting a global move toward more precise and controllable artificial intelligence.

360 Tech Engineering
360 Tech Engineering
360 Tech Engineering
Six 360 AI Institute Papers Accepted at 2026 Top AI Conferences Signal a Shift Toward Precise, Controllable AI

During the first half of 2026, the 360 Artificial Intelligence Institute had six research papers accepted at the four premier AI conferences—ICLR 2026, CVPR 2026, ICML 2026 and ECCV 2026—spanning multimodal understanding, intelligent agents, and image generation/editing. The overseas AI community has highlighted these works as part of a broader trend: AI is moving from raw capability demonstration toward precise, controllable applications.

Recent data show that Chinese mainland papers accounted for 43.7% of ICLR 2026 submissions, surpassing the United States for the first time, and the Stanford AI Index reports a narrowing performance gap of top US‑China models to 2.7%. As foundational models become more capable, competition shifts from building larger models to improving understanding, execution and real‑world applicability.

From "Seeing" to "Understanding": FG‑CLIP 2 (ICML 2026)

FG‑CLIP 2 tackles fine‑grained visual‑language alignment, enabling models to capture detailed correspondences between image regions and textual descriptions. The paper demonstrates that, unlike earlier models that struggle with subtle distinctions (e.g., a blue‑sofa cat versus a blue‑sofa dog), the enhanced alignment improves both overall scene comprehension and localized detail extraction, boosting performance in visual search and image understanding tasks.

From "Seeing" to "Understanding": AML (ICLR 2026)

AML focuses on agent goal‑location in realistic environments. As AI agents enter office‑automation and software‑operation scenarios, they must not only parse user commands but also accurately locate the target object amid complex layouts, visual clutter, or occlusions. By optimizing the linkage between visual regions and language instructions, AML improves agents' ability to identify the intended target, supporting more reliable interactive applications.

From "Generation" to "Precise Control": MoSA (ECCV 2026)

MoSA proposes a self‑supervised learning method that leverages video motion cues to let visual models autonomously learn object boundaries without human annotations. Using over 10,000 hours of unlabeled video, the approach automatically generates more than 21 million training labels, eliminating the need for costly manual labeling and opening a path toward large‑scale, low‑cost AI training.

From "Generation" to "Precise Control": RevealLayer (ICML 2026)

RevealLayer introduces a layered image‑understanding technique that decomposes an image into editable layers similar to design software, combined with fine‑grained control mechanisms. This enables precise edits of specific regions (e.g., changing a person’s appearance) without regenerating the entire image, advancing generative AI toward editable and controllable outputs for design, marketing and content creation.

From "Effect Boost" to "Efficiency Optimization": RefTON (CVPR 2026)

RefTON addresses virtual try‑on by guiding generation with reference images, reducing the need for auxiliary inputs such as pose data while preserving clothing texture, material and detail, thereby improving realism in virtual fitting applications.

From "Effect Boost" to "Efficiency Optimization": NAMI (CVPR 2026)

NAMI tackles high‑resolution image generation efficiency. By redesigning the model architecture, it maintains image quality while cutting inference cost, achieving roughly a 64% speed‑up for 1024×1024 resolution generation, which helps bring high‑quality generative AI to practical scenarios.

Beyond the papers, the institute has begun transferring technology to products: FG‑CLIP 2 powers intelligent search in 360 Cloud Disk and Yifang Cloud; RevealLayer’s methods are deployed on the 360 SaaS platform for smarter content processing. The overall narrative underscores that AI must evolve from merely generative capability to precise understanding, reliable execution and scalable deployment to create real industrial value.

Image
Image
Image
Image
Image
Image
Image
Image
Image
Image
Image
Image
Image
Image
Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

MultimodalAI researchagentsGenerative AIvision-languagetop conferencesprecision control
360 Tech Engineering
Written by

360 Tech Engineering

Official tech channel of 360, building the most professional technology aggregation platform for the brand.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.