Tagged articles

Benchmarks

11 articles · Page 1 of 1
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 16, 2026 · Artificial Intelligence

AI Solves Tests But Can't Self-Evolve: ByteDance Seed's Three RSI Benchmarks

ByteDance Seed and TokenWave introduce ASPIRE, S³Gym, and HarnessDev — three benchmarks that test whether AI agents can autonomously select learning goals, distill experience into improved decisions, and persistently upgrade their own execution systems without human-provided verification.

AI agentsASPIREBenchmarks
0 likes · 14 min read
AI Solves Tests But Can't Self-Evolve: ByteDance Seed's Three RSI Benchmarks
Machine Learning Algorithms & Natural Language Processing
Machine Learning Algorithms & Natural Language Processing
Sep 12, 2026 · Artificial Intelligence

How Robots Turn World Representations into Action: Jiajun Wu's ECCV 2026 Insights

Stanford professor Jiajun Wu's ECCV 2026 talk explores how structured world representations enable robots to act in complex physical environments, covering compositional skill learning, neural kinematics for zero-shot object interaction, video diffusion models for generating free demonstrations, and new benchmarks for evaluating embodied reasoning.

BenchmarksECCV 2026Embodied AI
0 likes · 35 min read
How Robots Turn World Representations into Action: Jiajun Wu's ECCV 2026 Insights
Machine Heart
Machine Heart
Sep 11, 2026 · Artificial Intelligence

PhyAgentOS v1.0.0: Executable, Verifiable, Evolvable Harness for Physical Agents

PhyAgentOS v1.0.0 introduces an open-source harness that unifies heterogeneous robots, verifies task outcomes with evidence, enables bounded recovery from failures, and evolves skills through verified experience, achieving 98.6% success on LIBERO and measurable gains on CALVIN and RoboCasa benchmarks.

BenchmarksEmbodied AIPhyAgentOS
0 likes · 17 min read
PhyAgentOS v1.0.0: Executable, Verifiable, Evolvable Harness for Physical Agents
Data Party THU
Data Party THU
Jun 30, 2026 · Artificial Intelligence

Large-Scale Sign Language Datasets: Resources, Benchmarks, and Annotation Standards

This ACL 2026 survey systematically reviews over 120 publicly available sign‑language datasets covering 35 languages, analyzes their modalities, annotation inconsistencies, and benchmark limitations, and proposes a 24‑field datasheet to promote reproducible and comparable AI research in sign language recognition, translation, and generation.

AI researchBenchmarksMultimodal
0 likes · 15 min read
Large-Scale Sign Language Datasets: Resources, Benchmarks, and Annotation Standards
Data Party THU
Data Party THU
Aug 23, 2025 · Industry Insights

How Apache IoTDB Dominated Benchmarks and Powered Industry 2023‑2025

The article summarizes the 2025 Time Series Database Innovation Conference where Apache IoTDB’s evolution, technical breakthroughs, benchmark leadership, open‑source community growth, and real‑world industrial deployments from aerospace to oil‑gas are detailed, highlighting the upcoming IoTDB 2.0 vision.

Apache IoTDBBenchmarksDB+AI
0 likes · 11 min read
How Apache IoTDB Dominated Benchmarks and Powered Industry 2023‑2025
AIWalker
AIWalker
May 11, 2025 · Artificial Intelligence

Unified Multimodal Understanding and Generation: A 30K‑Word Survey of Recent Advances

This comprehensive survey reviews the rapid progress of multimodal understanding and text‑to‑image generation models, categorises existing unified architectures into diffusion‑based, autoregressive, and hybrid paradigms, analyses their tokenisation strategies, datasets and benchmarks, and highlights current challenges and future research directions.

Benchmarksautoregressive modelsdatasets
0 likes · 64 min read
Unified Multimodal Understanding and Generation: A 30K‑Word Survey of Recent Advances
Architects' Tech Alliance
Architects' Tech Alliance
Jul 28, 2024 · Industry Insights

What Makes AMD’s Zen 5 Ryzen 9000 CPUs a Game‑Changer? Deep Dive into Architecture and Benchmarks

AMD’s Zen 5‑based Ryzen 9000 series and Ryzen AI 300 processors bring a 16% IPC boost, new NPU capabilities, and significant power‑efficiency gains, with detailed benchmark comparisons against Intel, Apple, and Qualcomm that reveal competitive advantages in both productivity and gaming workloads.

AMDBenchmarksCPU architecture
0 likes · 28 min read
What Makes AMD’s Zen 5 Ryzen 9000 CPUs a Game‑Changer? Deep Dive into Architecture and Benchmarks
Architects' Tech Alliance
Architects' Tech Alliance
May 18, 2018 · Industry Insights

Beyond Linpack: How HPCG, Graph500, and IO‑500 Redefine Supercomputer Rankings

This article examines why the traditional Linpack‑based TOP500 list is being complemented by newer benchmarks such as HPCG, Graph500, Green Graph 500 and IO‑500, explains their methodologies, presents the 2017 ranking results for major Chinese supercomputers, and reviews a wide range of application and micro‑benchmarks used to evaluate HPC system performance.

BenchmarksGraph500HPC
0 likes · 11 min read
Beyond Linpack: How HPCG, Graph500, and IO‑500 Redefine Supercomputer Rankings
Architects' Tech Alliance
Architects' Tech Alliance
Jul 4, 2017 · Industry Insights

Beyond Linpack: How HPCG and Graph500 Redefine Supercomputer Rankings

The article examines the 2017 TOP500, Green500, HPCG, Graph500 and Green Graph 500 rankings, explains why Linpack is becoming insufficient, compares benchmark methodologies, and introduces major application and micro‑benchmarks that illustrate the evolving performance metrics of modern supercomputers.

BenchmarksGraph500HPC
0 likes · 10 min read
Beyond Linpack: How HPCG and Graph500 Redefine Supercomputer Rankings