Tian Yuandong: Co-Authoring a Paper with GPT-5 Made Me Realize My Job Could Vanish in 5 Years

Meta FAIR veteran and Recursive Superintelligence co-founder Tian Yuandong recounts his journey from AlexNet debates to Go AI, Coconut, and grokking research, revealing how co-authoring a paper with GPT-5 boosted efficiency 6-10x and led him to predict AI will replace researchers within five years, while arguing humans must still initiate and verify ideas.

Machine Heart
Machine Heart
Machine Heart
Tian Yuandong: Co-Authoring a Paper with GPT-5 Made Me Realize My Job Could Vanish in 5 Years

Tian Yuandong, a near-decade veteran of Meta FAIR and co-founder of Recursive Superintelligence, joined The Information Bottleneck podcast to discuss his career trajectory and the shifting role of AI researchers. He began with his CMU PhD in non-convex optimization for computer vision, where hierarchical optimization mirrored the layered processing of convolutional networks. When AlexNet appeared in 2012, he recognized the connection but faced skepticism in Smith Hall, where most colleagues dismissed the result as luck.

At FAIR, Tian launched the DarkForest Go project in 2015, using convolutional networks to beat top amateurs. AlphaGo's subsequent victory over professionals surprised him not because of the method — convolutional nets capture board patterns humans rely on — but because it actually beat pros. He identified two gaps: DarkForest lacked a value network, and he had almost no reinforcement learning background, prompting his later deep dive into RL. In 2018, the team reproduced AlphaZero as OpenGo, achieving the first single-V100-GPU win against professional players in official matches.

Tian argues that in reinforcement learning, action-space and state-space design matter more than algorithms. He cites neural architecture search: if the agent searches kernel sizes first, all branches perform similarly and the value signal is useless; reorganizing the action space to let the agent "realize" depth importance creates discriminative subspaces. This insight led to Latent Action Monte Carlo Tree Search (LA-MCTS), which automatically finds high-discrimination action spaces.

On the Coconut paper (continuous latent-space chain-of-thought), Tian explains the idea came from introspection: he thinks without language, then translates thoughts for communication. Theoretical derivation showed latent representations can superpose multiple ideas, offering advantages on graph traversal. He gives two reasons frontier models haven't adopted continuous CoT: (1) human reasoning is recorded in language, not latent space; (2) latent representations may be highly personalized, and current models learn less efficiently than humans, so language serves as a proxy. He agrees with the host that this resembles the regularization challenges of JEPA-style self-supervision, and suggests gradient descent may be stuck in a meta-level local optimum, preventing discovery of better representations — a gap interpretability research could bridge.

The turning point came with his final solo-authored paper at Meta, studying the grokking phenomenon (networks memorize then suddenly generalize). He used GPT-5 as a collaborator: asking questions, assigning theorem proofs, checking proofs, and brainstorming. Efficiency rose 6-10x (a conservative estimate, since he still wrote code himself). This convinced him that researchers will be forced to think at a meta level rather than craft individual hypotheses, and that engineering leverage has surged — AI coding agents now let experienced researchers implement ideas rapidly.

Yet Tian insists humans remain essential for two roles: initiating (deciding which problems are worth solving) and verifying (judging whether AI proposals are sound). He notes recent agent-driven post-training and Kaggle attempts expose narrow idea spaces and metric overfitting, but believes providing more context can help models suggest more than hyperparameter tweaks. Ultimately, breaking out of routine work requires new architectures that capture new signals from little data — "a major leap."

On why alternative architectures fail at scale, Tian blames inductive bias: small data demands structural assumptions that help initially, but at massive scale the best approach is letting the model discover patterns itself. He cites ViT underperforming CNNs on ImageNet but surpassing them at hundreds of millions of images. Validation must start small, but with mechanistic understanding of what scales; otherwise only the compute-rich win.

His long-term goal is human-level learning efficiency: brains use ~20-30 watts and minimal data to understand quickly. He sees reinforcement learning and self-improvement as the fastest path. On recursive self-improvement regulation, he calls boundaries blurry — collaborating with a coding agent or co-authoring a paper already constitutes self-improvement — and pledges Recursive will open-source its research (kernels, training recipes) to avoid elite capture. He is optimistic open-source models will catch up (citing Kimi as already capable), and once they cross the usability threshold they become commoditized, letting everyone choose their preferred model.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

open-source AIreinforcement learninggrokkingrecursive self-improvementAI research automationarchitecture scalingCoconutGPT-5 collaboration
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.