Tagged articles

CyberGym

4 articles · Page 1 of 1
Machine Heart
Machine Heart
Jul 29, 2026 · Information Security

Chinese AI Beats OpenAI and Anthropic with 86.3% Success on CyberGym

Sangfor’s security‑focused AI, built on the domestic GLM‑5.2 model, completed 1,301 of 1,507 real‑world vulnerability tasks in the CyberGym benchmark, achieving an 86.3% success rate that places it among the global top‑four and demonstrates how evidence‑governed multi‑agent systems can turn model capabilities into verifiable security outcomes.

AI securityCyberGymEvidence governance
0 likes · 10 min read
Chinese AI Beats OpenAI and Anthropic with 86.3% Success on CyberGym
Black & White Path
Black & White Path
Jul 2, 2026 · Information Security

China’s Mysterious AI Security Team “MopMonk” Shocks the Industry with a 73% Success Rate

A previously unknown Chinese AI security group called MopMonk, operating without a website or corporate backing, posted a GitHub report that achieved a 73.1% vulnerability‑exploitation success rate, ranked seventh globally in the UC Berkeley‑run CyberGym benchmark, and demonstrated novel memory‑based multi‑agent techniques that signal China’s rising AI security prowess.

AI securityCyberGymMiniMax M3
0 likes · 9 min read
China’s Mysterious AI Security Team “MopMonk” Shocks the Industry with a 73% Success Rate
Black & White Path
Black & White Path
Jun 24, 2026 · Information Security

OpenAI’s GPT‑5.5‑Cyber Beats Mythos with 85.6% on CyberGym

OpenAI’s new GPT‑5.5‑Cyber model outperforms Anthropic’s Mythos on multiple security benchmarks, achieving 85.6% on CyberGym and 39.5% on ExploitGym, while the accompanying Daybreak initiative introduces the Codex Security plugin, Patch the Planet programme, and trusted‑access collaborations, prompting a shift in defensive priorities toward rapid patching.

AI securityCodex SecurityCyberGym
0 likes · 7 min read
OpenAI’s GPT‑5.5‑Cyber Beats Mythos with 85.6% on CyberGym
Machine Heart
Machine Heart
Jun 23, 2026 · Artificial Intelligence

How GPT‑5.5‑Cyber Beats Mythos 5 in CyberGym Benchmarks

OpenAI’s new GPT‑5.5‑Cyber model achieves a top‑of‑the‑line 85.6% score on CyberGym—surpassing both the prior GPT‑5.5 (81.8%) and Anthropic’s Mythos 5 (83.8%)—while also delivering broader security tools such as Codex Security, the Patch the Planet initiative, and a partner program for trusted access.

AI securityCodex SecurityCyberGym
0 likes · 12 min read
How GPT‑5.5‑Cyber Beats Mythos 5 in CyberGym Benchmarks