Tagged articles

Scientific Reasoning

3 articles · Page 1 of 1
Tencent Technical Engineering
Tencent Technical Engineering
Aug 28, 2026 · Artificial Intelligence

Tencent Hunyuan Hy4 Preview: 770B Open-Source Model Claims Top Tier

Tencent releases Hunyuan Hy4 preview, a 770B parameter mixture-of-experts model with 49B activated parameters and 1M context length, achieving top-tier open-source performance across coding, office, gaming, and scientific tasks, with internal blind tests showing 2.99/4 score surpassing GLM 5.3 and Kimi K3, plus 31.8% inference throughput gains via self-optimization.

HunyuanHy4LLM
0 likes · 6 min read
Tencent Hunyuan Hy4 Preview: 770B Open-Source Model Claims Top Tier
Machine Heart
Machine Heart
May 19, 2026 · Artificial Intelligence

100k‑Token Natural‑Language Reasoning Enables a 30B‑A3B Model to Reach Olympiad Gold Level

A 30B‑A3B model, trained with reverse‑perplexity supervised fine‑tuning, two‑stage reinforcement learning, and a multi‑round generate‑verify‑revise inference loop, achieves gold‑medal performance on IMO, USAMO and IPhO contests using over 100 k token natural‑language reasoning without external tools.

30B-A3BNatural Language ProcessingReinforcement Learning
0 likes · 11 min read
100k‑Token Natural‑Language Reasoning Enables a 30B‑A3B Model to Reach Olympiad Gold Level
HyperAI Super Neural
HyperAI Super Neural
Dec 18, 2025 · Artificial Intelligence

GPT-5 Leads as OpenAI Unveils FrontierScience: Dual‑Track Reasoning and Research Benchmark

OpenAI's FrontierScience benchmark, released on Dec 16, 2025, evaluates expert‑level scientific reasoning and research tasks, showing GPT‑5.2 scoring 25% on Olympiad and 77% on Research, outperforming other models while highlighting strengths in closed‑form problems and gaps in open‑ended research tasks.

AI evaluationBenchmarkFrontierScience
0 likes · 10 min read
GPT-5 Leads as OpenAI Unveils FrontierScience: Dual‑Track Reasoning and Research Benchmark