Tagged articles

self‑play reinforcement learning

2 articles · Page 1 of 1
Black & White Path
Black & White Path
Aug 13, 2026 · Information Security

How OpenAI’s GPT‑Red AI Red‑Team Automates Attacks in Four Steps, Outpacing Human Experts

OpenAI’s GPT‑Red model automates red‑team style prompt‑injection attacks through a four‑stage loop—goal setting, attack generation, response observation, and iterative refinement—demonstrating six‑fold safety gains over previous models and surpassing manual red‑team capabilities across multiple real‑world case studies.

AI securityGPT-RedLarge Language Models
0 likes · 29 min read
How OpenAI’s GPT‑Red AI Red‑Team Automates Attacks in Four Steps, Outpacing Human Experts
Baobao Algorithm Notes
Baobao Algorithm Notes
Sep 18, 2024 · Artificial Intelligence

How OpenAI’s o1 Uses Self‑Play RL to Achieve Breakthrough Reasoning

This article provides an in‑depth technical analysis of OpenAI’s new multimodal model o1, explaining its self‑play reinforcement‑learning pipeline, novel train‑time and test‑time scaling laws, inference‑time thinking process, and possible architectural variants, while also discussing broader implications for large‑language‑model research.

OpenAI o1Reward Modelinference thinking
0 likes · 37 min read
How OpenAI’s o1 Uses Self‑Play RL to Achieve Breakthrough Reasoning