Tagged articles

Multimodal Evaluation

7 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Oct 2, 2026 · Artificial Intelligence

PPTBench Benchmark Exposes Visual Coding Limits: GPT-6 Astra Leads but 20% Fail Semantic Checks

Einsia AI's PPTBench evaluates coding agents on reconstructing 500 scientific flowcharts into editable PPTX slides using a three-stage Agentic Judge; GPT-6 Astra High scores 77.34 with 80.8% passing semantic and rendering checks, yet semantic understanding remains the primary bottleneck across all models.

AI agentsAgentic JudgeBenchmark
0 likes · 13 min read
PPTBench Benchmark Exposes Visual Coding Limits: GPT-6 Astra Leads but 20% Fail Semantic Checks
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 30, 2026 · Artificial Intelligence

PPTBench: Can Coding Agents Accurately Reconstruct Flowcharts into Editable Slides?

Einsia AI's PPTBench benchmark evaluates coding agents on reconstructing 500 scientific flowcharts into editable PowerPoint slides, revealing that even top models like GPT-6 Astra fail 20% of semantic and rendering tests, with visual self-inspection correlating strongly (r=0.88) with success.

AI agentsAgentic JudgeFlowchart Understanding
0 likes · 14 min read
PPTBench: Can Coding Agents Accurately Reconstruct Flowcharts into Editable Slides?
Bilibili Tech
Bilibili Tech
Jul 1, 2026 · Artificial Intelligence

FATE Series (SABER & CASTER) Debuts at ACL 2026: Advanced LLM Reasoning

At ACL 2026 in San Diego, Bilibili’s tech team introduced the FATE series—SABER, which reduces overthinking in LLMs with a token‑budgeted switchable training, and CASTER, a community‑aware evaluation system built on Social‑CoT and the MEDEA framework that outperforms GPT‑5.2 and Claude‑4.5‑Opus on the new CASTER‑Bench, while also promoting the B‑UP talent recruitment program.

BenchmarkLLMMultimodal Evaluation
0 likes · 10 min read
FATE Series (SABER & CASTER) Debuts at ACL 2026: Advanced LLM Reasoning
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
May 29, 2026 · Artificial Intelligence

WBench: 20 Cutting‑Edge World Models Face a Comprehensive Interactive Benchmark

WBench, a new benchmark created by Meituan LongCat and Fudan University, evaluates 20 state‑of‑the‑art video and world‑model systems across 289 test cases and 1,058 interaction rounds, measuring video quality, setting adherence, interaction fidelity, consistency and physical compliance, and reveals that no model yet excels in all five dimensions.

Interactive BenchmarkMultimodal EvaluationWBench
0 likes · 10 min read
WBench: 20 Cutting‑Edge World Models Face a Comprehensive Interactive Benchmark
SuanNi
SuanNi
Mar 2, 2026 · Artificial Intelligence

Why Leading AI Models Flunk the New ‘Humanity’s Last Exam’ Benchmark

The newly released Humanity’s Last Exam (HLE) benchmark, featuring 2,500 rigorously crafted multimodal questions across more than 100 disciplines, exposes the severe shortcomings of leading AI models, whose accuracy stays below 50% and shows alarming calibration errors, highlighting the urgent need for deeper AI evaluation.

Artificial IntelligenceHumanity's Last ExamMultimodal Evaluation
0 likes · 13 min read
Why Leading AI Models Flunk the New ‘Humanity’s Last Exam’ Benchmark
Tencent Technical Engineering
Tencent Technical Engineering
Jun 30, 2025 · Artificial Intelligence

How iMatch Won CVPR2025 NTIRE Image-Text Alignment: Techniques & Benchmarks

The IH‑VQA team’s iMatch solution clinched the CVPR2025 NTIRE Image‑Text Alignment champion by introducing dual‑model fusion, pseudo‑label data augmentation, Q‑Align probability mapping, and visual augmentations, and the paper also presents a comprehensive iMatch benchmark evaluating 23 state‑of‑the‑art text‑to‑image models across multiple resolutions.

AI quality assessmentCVPR2025Multimodal Evaluation
0 likes · 15 min read
How iMatch Won CVPR2025 NTIRE Image-Text Alignment: Techniques & Benchmarks
Sohu Tech Products
Sohu Tech Products
Jul 31, 2024 · Artificial Intelligence

MMEvalPro: A Trustworthy Benchmark for Evaluating Multimodal Large Models

MMEvalPro, a new benchmark created by researchers from Peking University, Chinese Academy of Medical Sciences, CUHK and Alibaba, augments existing multimodal datasets with perception and knowledge questions and introduces a Genuine Accuracy metric, revealing that top multimodal models still lag far behind humans and exposing shortcut‑driven performance on prior tests.

BenchmarkLarge Language ModelsMMEvalPro
0 likes · 11 min read
MMEvalPro: A Trustworthy Benchmark for Evaluating Multimodal Large Models