AlphaGo Co-Author Warns: LLMs Lack True Reasoning Despite Move 37 Breakthrough

AlphaGo co-author Thore Graepel argues that LLMs only scale intuitive pattern matching without a separate deliberative search mechanism, making them fundamentally different from AlphaGo's hybrid architecture and unreliable for high-stakes reasoning tasks.

Machine Heart
Machine Heart
Machine Heart
AlphaGo Co-Author Warns: LLMs Lack True Reasoning Despite Move 37 Breakthrough

AlphaGo's Move 37: Intuition vs. Search

In the second game of the 2016 match against Lee Sedol, AlphaGo played move 37 on the fifth line — a move that professional commentators initially suspected was a bug because human professionals would almost never play it (probability ~1 in 10,000). Thore Graepel, a core AlphaGo author, explains that this move was not a flash of intuition but the product of explicit search. AlphaGo combines two systems: a policy network (System 1) trained on human games to predict likely moves, and a Monte Carlo tree search (System 2) that builds a game tree of thousands of branches to evaluate each candidate move's long-term consequences. The policy network alone would never have selected move 37; the search mechanism discovered it led to a better terminal position.

Graepel contrasts this with Deep Blue (1997), which relied on hand-crafted rules and brute-force evaluation of 200 million positions per second to a depth of 6–8 plies. Go's complexity — where a stone's value depends on developments dozens of moves later — makes pure brute force infeasible; even a supercomputer would need billions of years to explore a fraction of the game tree. AlphaGo succeeded by learning "intuition" (value/policy networks) to guide search, not by compute alone.

LLMs as a Scaled System 1

Large language models operate by repeatedly predicting the next token — a process Graepel equates to Kahneman's System 1: fast, associative, and pattern-completing across vast domains. Chain-of-thought prompting improves performance by forcing the model to generate intermediate steps, but Graepel argues these steps are still produced by the same next-token prediction process, merely extended in time. They do not introduce an independent reasoning mechanism analogous to AlphaGo's search.

Three Shortcomings of LLM "Reasoning"

No explicit, inspectable cognitive state: The model does not maintain a structured record of hypotheses, confidence levels, evidence weighed, or open questions that can be systematically updated as new information arrives.

No separation of knowledge and reasoning: Facts and inference procedures are entangled in the network weights; there is no distinct, explicit belief system that can be examined or manipulated independently.

Post-hoc rationalization: Research shows models often fabricate plausible-sounding reasoning traces after the fact; the reported chain of thought may not reflect the actual computational path to the answer.

Why Reasoning Fidelity Matters

In high-stakes domains — medical diagnosis, engineering, scientific research — it is essential to audit how a conclusion was reached: which step failed, which evidence was faulty, which assumption was wrong. A system that only produces convincing post-hoc stories cannot be trusted or held accountable.

Graepel's Proposed Architecture

Graepel left DeepMind to build a general reasoning system that maintains an explicit cognitive state — a structured record of confirmed beliefs, doubts, excluded options, and open questions. Reasoning becomes a sequence of actions that modify this state: deriving conclusions, decomposing problems, and crucially, deciding what question to ask, what computation to run, or what experiment to perform next. An independent "judge" evaluates each action by how much uncertainty it resolves, updating beliefs only when evidence supports them. LLMs and other neural models serve as components: proposing solution paths, interacting with tools via APIs, and assessing whether evidence supports a claim. Graepel describes this as "the scientific method on steroids" — producing knowledge that withstands scrutiny.

Debate and Criticism

The article sparked discussion on X. University of Toronto professor David Duvenaud noted Graepel's three criticisms also apply to unaided human reasoning. Graepel replied that humans reason poorly without external structure (paper, pencil, scientific method); the key is structuring the process to avoid cognitive bias. Duvenaud inferred that equipping LLMs with similar tools might satisfy Graepel's definition — which Graepel implicitly confirmed. Other commenters questioned whether this amounts to "current model + external system" rather than a stronger internal reasoning model. Graepel argued that frontier labs bet on scaling, but auditable reasoning cannot emerge from scaling alone; it requires a different architecture. A further comment suggested natural language itself is the wrong substrate for reliable reasoning, raising the open question of whether a formal language for real-world reasoning could be designed.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LLMreasoningAI architectureAlphaGoSystem 1System 2MIT Technology ReviewThore Graepel
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.