Tagged articles

Process Evaluation

2 articles · Page 1 of 1
Machine Heart
Machine Heart
Jul 6, 2026 · Artificial Intelligence

Evaluating Multi-Agent LLM Systems: Rethinking the Orchestrator’s Role

The paper reveals that failures in LLM‑driven multi‑agent systems often stem from the Orchestrator’s loss of control, introduces an entropy‑dynamics framework to measure scheduling entropy, and proposes Inverse Workflow Generation for detailed process evaluation, shifting focus from agent strength to orchestration stability.

Entropy DynamicsICML 2026LLM
0 likes · 11 min read
Evaluating Multi-Agent LLM Systems: Rethinking the Orchestrator’s Role
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Mar 13, 2026 · Artificial Intelligence

Can Multimodal LLMs Beat Humans in Real Web Search? GPT‑5.2 Scores Only 36% on New BrowseComp‑V3 Benchmark

A new multimodal browsing benchmark, BrowseComp‑V3, reveals that human experts achieve a 68.03% success rate while the strongest closed‑source model, GPT‑5.2, manages just 36.17%, highlighting current limitations in deep web‑scale visual‑text reasoning and the critical role of tool‑augmented agents.

GPT-5.2OmniSeekerProcess Evaluation
0 likes · 12 min read
Can Multimodal LLMs Beat Humans in Real Web Search? GPT‑5.2 Scores Only 36% on New BrowseComp‑V3 Benchmark