Tagged articles

VLM

15 articles · Page 1 of 1
JD Retail Technology
JD Retail Technology
Jul 13, 2026 · Artificial Intelligence

Inside JD’s Oxygen AIIC: An Industrial‑Scale LLM/VLM‑Powered Product Knowledge Platform for Billions of SKUs

JD’s Oxygen AIIC combines human‑in‑the‑loop ontology engineering, a semantic search‑then‑discrimination pipeline, and a self‑evolving multi‑task LLM/VLM model to produce high‑quality product knowledge for over a hundred thousand categories and billions of daily SKU updates, boosting search coverage to 80%, attribute auto‑fill to over 80%, cutting quality issues by 37% and raising click‑through by 9% while achieving 94.2% precision and 82.8% recall.

JD.comKnowledge GraphLLM
0 likes · 21 min read
Inside JD’s Oxygen AIIC: An Industrial‑Scale LLM/VLM‑Powered Product Knowledge Platform for Billions of SKUs
Machine Heart
Machine Heart
Jul 8, 2026 · Artificial Intelligence

How DOPD Overcomes the Privilege Illusion to Boost Online Policy Distillation

The DOPD paper introduces an advantage‑aware dual distillation framework that eliminates the privilege illusion, dynamically selects token‑wise strategies, and delivers up to 7.5‑point gains on LLM benchmarks while closing 89.8% of the teacher‑student gap and showing strong robustness across model sizes.

LLMVLMdual on-policy distillation
0 likes · 9 min read
How DOPD Overcomes the Privilege Illusion to Boost Online Policy Distillation
DeWu Technology
DeWu Technology
Jul 1, 2026 · Artificial Intelligence

AI UITester: The New AI‑Native Paradigm for Visual UI Automation Testing

The article analyzes the limitations of traditional UI automation, introduces the AI‑Native ai_uitester pipeline that converts test‑case data with LLM enhancement, implements AI‑driven debugging and self‑healing, and unifies cross‑platform execution through a VLM‑based engine, backed by real‑world metrics.

AI testingLLMPrompt Engineering
0 likes · 17 min read
AI UITester: The New AI‑Native Paradigm for Visual UI Automation Testing
Machine Heart
Machine Heart
Jun 23, 2026 · Artificial Intelligence

Doubao Model 2.1 Launch: Production‑Grade End‑to‑End Coding and Multi‑Agent Breakthrough

Doubao's Model 2.1, unveiled at the Force conference, pushes daily token usage past 180 trillion, captures 49.5% of China's public‑cloud MaaS market, tops code and agent benchmarks, delivers repository‑level coding, advanced multi‑modal reasoning, and introduces cost‑effective Pro and Turbo variants with a new Deep Think inference mode.

AI benchmarkingDoubaoLLM
0 likes · 11 min read
Doubao Model 2.1 Launch: Production‑Grade End‑to‑End Coding and Multi‑Agent Breakthrough
Qunhe Technology Quality Tech
Qunhe Technology Quality Tech
Jun 23, 2026 · Artificial Intelligence

Why Pixel Diff Failed and How VLM Fine‑Tuning Became the Eyes of UI Automation

Traditional pixel‑by‑pixel UI comparison breaks on complex CAD drawings due to semantic changes, so a team built a visual‑language‑model fine‑tuning pipeline that turns failure cases into training data, achieves ~95% AI accuracy, improves regression efficiency by over 40%, and now powers hundreds of daily automation tests.

AI monitoringFine-tuningModel Evaluation
0 likes · 12 min read
Why Pixel Diff Failed and How VLM Fine‑Tuning Became the Eyes of UI Automation
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Jun 18, 2026 · Artificial Intelligence

UniRL: Tencent Hunyuan’s Open‑Source Framework Unifying Multimodal RL Training

UniRL is an open‑source, distributed reinforcement‑learning post‑training framework that consolidates fragmented pipelines for image, video, and language‑vision models, offering a unified rollout‑reward‑advantage‑train‑sync contract, extensive model support, built‑in algorithms, and multi‑modal reward components to lower engineering barriers in AIGC research.

Distributed TrainingLLMMultimodal RL
0 likes · 10 min read
UniRL: Tencent Hunyuan’s Open‑Source Framework Unifying Multimodal RL Training
Machine Heart
Machine Heart
Jun 9, 2026 · Artificial Intelligence

Why Standard Vision‑Language Models + Scale Data Beat Specialized 3D Vision Designs (VLM³)

Meta’s VLM³ demonstrates that a plain vision‑language model, when trained on large‑scale data with simple camera‑focal‑length and pixel‑space normalization, matches or surpasses expert 3D vision models across monocular depth estimation, object‑level understanding, pixel‑matching and camera‑pose tasks, eliminating the need for task‑specific architectures, loss functions, data augmentations or regression formulations.

3D visionDepth EstimationMeta
0 likes · 6 min read
Why Standard Vision‑Language Models + Scale Data Beat Specialized 3D Vision Designs (VLM³)
Machine Heart
Machine Heart
May 31, 2026 · Artificial Intelligence

How a Near‑Invisible Image Can Make GPT‑5.4 and Claude Opus 4.6 Spread False Claims

Researchers from ETH Zurich show that tiny, human‑imperceptible perturbations to a single image can fool leading visual language models—including GPT‑5.4, Claude Opus 4.6, and Grok—into confidently delivering fabricated answers, enabling misinformation amplification, defamation, content‑filter evasion, and large‑scale AI authority laundering.

AI safetyClaude OpusGPT-5.4
0 likes · 7 min read
How a Near‑Invisible Image Can Make GPT‑5.4 and Claude Opus 4.6 Spread False Claims
AI Engineer Programming
AI Engineer Programming
May 9, 2026 · Artificial Intelligence

Why PDF Parsing Is Hard for RAG and Which Mainstream Solutions Work

The article examines the intrinsic challenges of extracting structured text from PDFs for Retrieval‑Augmented Generation—such as missing reading order, table reconstruction, font encoding, and scanned images—and compares lightweight libraries, AI‑enhanced frameworks, commercial APIs, and visual language models as practical solutions.

AI frameworksOCRPDF parsing
0 likes · 23 min read
Why PDF Parsing Is Hard for RAG and Which Mainstream Solutions Work
Old Zhang's AI Learning
Old Zhang's AI Learning
Jan 30, 2026 · Artificial Intelligence

PaddleOCR‑VL‑1.5: 0.9B Model Beats Billion‑Parameter OCR Models with 94.5% Accuracy

PaddleOCR‑VL‑1.5, the latest Baidu release, uses only 0.9 B parameters to achieve 94.5% accuracy on OmniDocBench v1.5, surpassing larger open‑source and commercial OCR models, while offering multi‑task, multi‑language support, lightweight deployment, and detailed performance benchmarks.

DeepSeek-OCRGPU inferenceMulti-language
0 likes · 9 min read
PaddleOCR‑VL‑1.5: 0.9B Model Beats Billion‑Parameter OCR Models with 94.5% Accuracy
Data Party THU
Data Party THU
Nov 5, 2025 · Artificial Intelligence

How VLM‑FO1 Turns Vision‑Language Models into Precise Perception Machines

VLM‑FO1 introduces a generate‑plus‑reference paradigm that replaces coordinate generation with region token referencing, adding plug‑in modules such as a proposal generator, a hybrid fine‑grained encoder, and a region‑language connector to give any pretrained visual language model accurate, fine‑grained perception while preserving its original capabilities.

AI researchMultimodalPlug-and-Play
0 likes · 15 min read
How VLM‑FO1 Turns Vision‑Language Models into Precise Perception Machines
AI Algorithm Path
AI Algorithm Path
Jul 20, 2025 · Artificial Intelligence

How to Build an Open‑Set Object Detection Workflow: A Comprehensive Guide

This article presents a step‑by‑step agentic object detection pipeline that combines open‑vocabulary detectors such as Grounding‑DINO with visual language models (GPT‑4o, o1) for concept extraction, critique, refinement, and validation, complete with code snippets, design rationale, and real‑world examples.

Grounding DINOOpen-Vocabulary DetectionPython
0 likes · 33 min read
How to Build an Open‑Set Object Detection Workflow: A Comprehensive Guide
Sohu Tech Products
Sohu Tech Products
Jan 8, 2025 · Artificial Intelligence

Multimodal RAG: Implementation Paths and Development Prospects

The talk outlines Multimodal RAG implementation routes, comparing OCR‑based object recognition, transformer encoder‑decoder encoding, and Visual Language Model processing, explains the ColPali late‑interaction method for multi‑dimensional vector matching, addresses scaling tensors with binarization and reranking, and recommends a hybrid long‑term strategy where VLM excels on abstract imagery while traditional OCR remains valuable.

ColPaliDocument processingMultimodal RAG
0 likes · 10 min read
Multimodal RAG: Implementation Paths and Development Prospects