Richard Sutton on Energy‑Efficient AI, Over‑Hyped Large Models, and Alignment

In a candid WAIC 2026 interview, reinforcement‑learning pioneer Richard Sutton discusses his new for‑profit Oak Lab, the quest for a 20‑watt trillion‑parameter model, his disappointment with recent AI trends, the notion of a “complete mind,” robot‑kindergarten experiments, and why he believes aligning AI to a single human value system is a dangerous illusion.

Machine Heart
Machine Heart
Machine Heart
Richard Sutton on Energy‑Efficient AI, Over‑Hyped Large Models, and Alignment

Meeting Richard Sutton at WAIC 2026

During WAIC 2026, reinforcement‑learning founder and 2024 Turing Award laureate Richard Sutton sat down for a group interview. He explained that he now runs two organizations: the for‑profit Oak Lab and the non‑profit Openmind Research Institute . Oak Lab is meant to complement his nonprofit work, allowing short‑term commercial projects to fund longer‑term open research.

Energy‑Efficient AI and the 20‑Watt Goal

Sutton said Oak Lab’s long‑term ambition is to build a trillion‑parameter model that learns and plans in real time while consuming only about 20 watts . He argued that current large‑language‑model (LLM) approaches waste energy and compute because digital architectures move data back and forth between storage and CPUs. He suggested that massive parallelism that keeps data “in place” would dramatically improve efficiency.

What a “Complete Mind” Means

The researcher emphasized his interest in a “ complete mind ” rather than a narrow tool. He distinguishes this from the popular “agent” concept, saying a complete mind would be a general intelligence that could use specialized tools, not merely a system optimized for a single task.

Looking Back on Eight Years of AI

When asked to evaluate the past eight years, Sutton expressed a “certain degree of disappointment.” He noted that after the AlphaGo/AlphaZero era, the field reverted to heavy reliance on human‑generated data instead of learning from experience. He sees a shift back toward agency, with researchers exploring agents, computer use, and AI‑driven mathematics. Sutton is optimistic about the next eight years, betting on an “experience‑based” era that could achieve recognition comparable to today’s LLMs.

“Robot Kindergarten” Project

Sutton described a Beijing‑based “Robot Kindergarten” initiative aimed at creating robots robust enough to learn through trial‑and‑error for weeks or months, similar to how infants learn to walk. He argued that conventional robots are built to follow precise engineer commands and break quickly when used for experimentation.

Learning Feels Good

He offered a mechanistic view of infant learning: babies explore, try actions, and monitor their own progress. When progress stalls, they switch to a new activity, maximizing overall learning speed. Sutton summed this up with the slogan “ learning feels good ,” suggesting that intrinsic reward signals could be designed to capture this feeling.

World Models vs. Reinforcement Learning

Sutton clarified terminology. He insists a “world model” should be understood as a “state‑transition” model that encodes knowledge, not merely a physics simulator. Likewise, he argues that reinforcement learning already includes imagination, planning, and model‑based components; the term “transition model” is the textbook name for this part.

Advice for Young Researchers

When asked what he would tell a 20‑year‑old immersed in ready‑made answers, Sutton replied: “Don’t get sucked into that system.” He recommends two habits: (1) return to first‑principles and verify conclusions with measurable evidence, and (2) keep a notebook and write a page each day to capture and refine ideas. He also stresses learning to harness ever‑stronger compute across domains, estimating that AI‑level intelligence may become mainstream around 2030‑2040.

Alignment as a Human Trick

On AI safety, Sutton challenged the premise of aligning AI to “human values,” arguing that no single set of human values exists. He likened alignment attempts to a dangerous illusion, suggesting that if we could force alignment on people, it would be an evil tool. He concludes that the world’s greatest asset is that “human beings cannot be fully aligned,” which, in his view, preserves peace.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

reinforcement learningAI safetyAI alignmentenergy efficiencyRichard Suttonexperience learningcomplete mindrobot kindergarten
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.