DeepSeek-V4-Flash-Vision-Exp Multimodal Evaluation: Strong Recognition but Over-Inference on Real-World Context
The author evaluates DeepSeek's new multimodal model across four visual reasoning challenges, finding excellent recognition and structured reasoning capabilities but a consistent tendency to hallucinate real-world details not present in images, a common limitation in vision-language models.
DeepSeek released DeepSeek-V4-Flash-Vision-Exp, its first experimental multimodal model. The author tests it via the DeepSeek Harness (DSH) webui across four challenges of increasing difficulty, focusing on multimodal reasoning — OCR, spatial understanding, counting, logical reasoning, temporal inference, and anomaly detection — rather than simple image recognition.
Challenge 1: Detective Desk Clues
根据图片内容,回答下面三个问题:
图片中有哪些人物?
主人今天有哪些安排?
根据所有线索,推断主人下午是否可能赶不上火车,并说明理由。The model quickly analyzes the scene from five angles and correctly links multiple elements (train ticket, supermarket receipt, phone time, dinner plan). However, two over-inferences appear:
Train ticket vs. receipt: The ticket shows 14:15 departure from Beijing; the receipt shows 17:56 supermarket purchase. The model treats the receipt as proof the owner was at the supermarket, ignoring that a family member could have shopped.
Phone time vs. location: Phone shows 18:42 and a 19:15 dinner with "Ale". The model assumes the owner is still local based solely on phone time, which does not prove physical presence.
These reveal gaps in uncertainty judgment and anti-hallucination capability.
Challenge 2: Supermarket Shelf
根据图片内容,回答下面三个问题:
哪个商品缺货?
货架上最贵商品是什么?
买这些商品需要多少钱?Results are strong: identification, counting, and calculation are accurate with no obvious hallucinations. The model excels at OCR, product statistics, and arithmetic on structured visual data.
Challenge 3: Station Information
根据图片内容,回答下面三个问题
哪辆车最适合乘坐?
从售票处到 B1 检票口怎么走?
一个带老人和行李的人,现在在候车大厅,要赶 09:05 的车,应该如何规划路线?The model reads train numbers, times, and gate codes correctly but over-completes the navigation:
It suggests "take accessible elevator, avoid stairs" and describes walking "through the hall center, then south to B1-B10". The image contains no elevator/stair locations, accessibility facilities, or cardinal directions — these are hallucinated real-world completions.
A rigorous answer would state only what the map shows: "B gate area is below the waiting hall; proceed to B1-B10."
Challenge 4: Parking Lot
根据图片内容,回答下面四个问题
A3 是否有车?
C1 EV 车位的白色车是否违规停车?
车辆 B5 应该如何驶出停车场?
找一个,靠近出口,非新能源区域,空闲的停车位置This test took 48 reasoning steps and over 10 minutes in DSH minimal mode. Performance is stable and better than Challenge 3, showing strength in rule-based spatial understanding. Yet hallucination persists: the model assumes an accessible elevator near the exit that is not depicted.
Overall Findings
Across all four tests, a consistent pattern emerges:
Strong visual recognition and structured reasoning, but a tendency to over-infer real-world common sense not present in the image. This is not unique to DeepSeek; it is a common weakness of current vision-language models.
Testing cost: approximately ¥2.5 total.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Sohu Tech Products
A knowledge-sharing platform for Sohu's technology products. As a leading Chinese internet brand with media, video, search, and gaming services and over 700 million users, Sohu continuously drives tech innovation and practice. We’ll share practical insights and tech news here.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
