Tagged articles

benchmark results

10 articles · Page 1 of 1
Amap Tech
Amap Tech
Jul 10, 2026 · Artificial Intelligence

ABot-M0.5: The First Unified World Action Model for Mobile Manipulation

ABot-M0.5 introduces a unified world action model that aligns video prediction, intermediate latent actions, and low‑level physical control for mobile manipulation, achieving state‑of‑the‑art long‑horizon success rates and fine‑grained precision across benchmarks such as RoboCasa365, RoboTwin, and LIBERO, while detailing novel architectural components and a three‑stage progressive training regime.

benchmark resultsdream forcingdual-level transformer
0 likes · 14 min read
ABot-M0.5: The First Unified World Action Model for Mobile Manipulation
Data Party THU
Data Party THU
Jun 21, 2026 · Artificial Intelligence

Lance: A Lightweight 3B Multimodal AI Model that Handles Vision, Video, Generation, and Editing

Lance, an open‑source 3‑billion‑parameter multimodal model from ByteDance, unifies image and video understanding, generation, and editing in a single architecture, achieves top scores on VBench (85.11), MVBench (62.0), GenEval (0.90) and GEdit‑Bench (7.30), and demonstrates emergent cross‑task generalization.

LanceMaPEbenchmark results
0 likes · 9 min read
Lance: A Lightweight 3B Multimodal AI Model that Handles Vision, Video, Generation, and Editing
Machine Heart
Machine Heart
May 24, 2026 · Artificial Intelligence

Inside the First Vision-Centric Parallel Thinking Framework for Vision-Language Models

The article introduces Visual Para-Thinker, the first parallel reasoning framework tailored for large‑scale vision‑language models, explains its block and scan visual path divisions, details the Path‑aware Attention and Learnable Parallel Rotary Position Embedding mechanisms, and presents experimental results showing significant gains on visual perception benchmarks.

LPRoPEPath-aware AttentionVision-Language Models
0 likes · 9 min read
Inside the First Vision-Centric Parallel Thinking Framework for Vision-Language Models
Tencent Technical Engineering
Tencent Technical Engineering
Apr 23, 2026 · Artificial Intelligence

Tencent Hunyuan Launches Hy3 Preview: Open‑Source Model Boosts Agent Performance

On April 23, Tencent released the open‑source Hy3 preview, a 295 B‑parameter hybrid expert model with 21 B active parameters and 256K context length, delivering substantial gains in complex reasoning, instruction following, code and agent tasks, achieving 40 % faster inference, lower costs, and strong benchmark results across Tencent’s AI products.

Hy3-previewTencent Hunyuanagent capabilities
0 likes · 9 min read
Tencent Hunyuan Launches Hy3 Preview: Open‑Source Model Boosts Agent Performance
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Apr 7, 2026 · Artificial Intelligence

Can AI Self‑Evolve? New Meta Research Redefines Agent Rules

A recent Meta‑led study introduces HyperAgents, a framework that merges task agents with meta‑agents to enable metacognitive self‑modification, showing significant gains on coding benchmarks, paper review, robotics reward design, and Olympiad‑level math grading, while also highlighting emerging safety risks as AI systems begin to rewrite their own improvement mechanisms.

Darwin Gödel MachineHyperagentsSelf‑Improvement
0 likes · 10 min read
Can AI Self‑Evolve? New Meta Research Redefines Agent Rules
AI Insight Log
AI Insight Log
Feb 18, 2026 · Artificial Intelligence

Claude Sonnet 4.6 Launches on Chinese New Year with Opus-Level Coding Power

Anthropic unveiled Claude Sonnet 4.6 on February 18, touting Opus-level coding ability, a 1 million-token context window, and unchanged pricing; benchmarks show a SWE-bench score of 79.6% (up from 77.2%), OSWorld 72.5% (vs 61.4%), and GPQA Diamond 89.9%, while industry leaders praise its reduced laziness, stronger instruction following, and strategic long-term planning.

AI CodingAnthropicClaude Sonnet 4.6
0 likes · 7 min read
Claude Sonnet 4.6 Launches on Chinese New Year with Opus-Level Coding Power
AI Frontier Lectures
AI Frontier Lectures
Nov 25, 2025 · Artificial Intelligence

How RoMa v2 Achieves Harder, Better, Faster, Denser Feature Matching

RoMa v2 introduces a two‑stage matching‑then‑refinement pipeline powered by DINOv3 features, custom CUDA kernels, and diverse training data, delivering state‑of‑the‑art accuracy, speed, and pixel‑level uncertainty estimation across a wide range of dense matching benchmarks.

DINOv3RoMa v2benchmark results
0 likes · 10 min read
How RoMa v2 Achieves Harder, Better, Faster, Denser Feature Matching
Bighead's Algorithm Notes
Bighead's Algorithm Notes
Oct 17, 2025 · Artificial Intelligence

Exploring MLLM4TS: A Universal Multimodal Framework for Time‑Series Analysis

This article reviews the MLLM4TS framework, which fuses visual representations of multivariate time series with large language models to address complex temporal dependencies, cross‑channel interactions, and task generalization, and demonstrates superior performance on classification, anomaly detection, forecasting, and few‑shot scenarios across multiple benchmarks.

Ablation StudyFew-shot LearningMultimodal LLM
0 likes · 11 min read
Exploring MLLM4TS: A Universal Multimodal Framework for Time‑Series Analysis
AI Frontier Lectures
AI Frontier Lectures
May 25, 2025 · Artificial Intelligence

Can Alternating Generation‑Reduction Make LLMs Think Faster? Introducing PENCIL

The paper presents PENCIL, a novel alternating generation‑and‑erasure reasoning paradigm that achieves optimal space‑time complexity for chain‑of‑thought tasks, dramatically improves accuracy and efficiency on hard SAT, QBF, and Einstein puzzle benchmarks, and is provably Turing‑complete.

Chain of ThoughtPencilbenchmark results
0 likes · 12 min read
Can Alternating Generation‑Reduction Make LLMs Think Faster? Introducing PENCIL
AIWalker
AIWalker
Apr 6, 2025 · Artificial Intelligence

NOVA: Redefining Autoregressive Visual Modeling Without Vector Quantization

NOVA introduces a highly efficient autoregressive video generation framework that eliminates vector quantization, combines frame‑by‑frame causal prediction with set‑by‑set spatial attention, and achieves state‑of‑the‑art quality on VBench and GenEval while offering strong zero‑shot generalization across text‑to‑image and text‑to‑video tasks.

Novaautoregressive video generationbenchmark results
0 likes · 14 min read
NOVA: Redefining Autoregressive Visual Modeling Without Vector Quantization