Karpathy: Replace Raw AI Text with Diagrams, Interactive Pages & Videos
Andrej Karpathy argues that AI now generates output faster than humans can comprehend, proposing four formats — simplified technical English, diagrams, interactive web pages, and explanatory videos — to make AI outputs verifiable, while warning that richer formats make quality control harder.
Andrej Karpathy recently posted a thread highlighting a growing mismatch: models can produce tens of thousands of lines of code, research reports, or complex designs in minutes, but humans still need to judge what the model did, whether the logic holds, and whether the result is usable. He proposes four ways to transform model output into forms that are easier to understand and verify.
1. Constrain text with ASD‑STE100
Standard models tend to use long sentences, abstract vocabulary, and vague phrasing. ASD‑STE100 (Simplified Technical English), originally developed for aircraft maintenance manuals, restricts vocabulary, sentence structure, and expression so that each sentence is clear, direct, and unambiguous. A prompt such as "Explain this in 80% ASD‑STE100 style" reduces linguistic noise and makes technical documents easier to audit.
2. Generate diagrams
Complex relationships — system architecture, call chains, causal links — are hard to convey in linear text. Asking the model to produce flowcharts, architecture diagrams, or concept maps lets readers see how nodes connect, how data flows, and where problems may arise. This is essentially what UML aimed to do, but previously required manual drawing; now an LLM can generate a diagram directly from the specific content at hand.
3. Generate interactive web pages
When a problem involves state, parameters, or dynamic changes, static text and images are insufficient. The model can output HTML with buttons, sliders, and animations that show how results change when parameters are adjusted or how a system behaves in different states. This turns "reading a conclusion" into "hands‑on verification." A community member noted that interactive pages let users tweak parameters and immediately see causal relationships; another suggested a "guess‑then‑see" mode where the result is hidden until the user predicts the outcome.
4. Generate explanatory videos
Algorithms, physical processes, and abstract concepts all have a temporal dimension. Video can walk through the steps with animation, narration, and highlights. Karpathy suggests asking the model to create 3Blue1Brown‑style explanatory videos. One user generated a 4‑minute video on recursive self‑improvement using Opus 5.5; the source HTML was 13.8 MB but rendered to 34 GB. Karpathy endorsed it. Another user produced a stochastic‑calculus video in a similar style.
The underlying shift: disposable software
These approaches share a common driver: the cost of "disposable software" — custom web pages, animations, or videos built for a single question — is approaching zero. Previously, creating a bespoke visualization for one user was uneconomical; now a model can generate it on demand, tailored to the user's knowledge level, and discard it after use. The takeaway: stop forcing yourself to read every word of raw model output; let the model re‑express it in a format that lowers the comprehension cost.
Warning: richer formats make quality control harder
Easier‑to‑consume formats also make errors harder to spot. A flawed paragraph can be scanned quickly; a polished video with animation, narration, and visual effects distributes attention, so mistakes hide inside attractive packaging. Commenters observed that a good‑looking erroneous video can be more persuasive than a flawed text. In interactive papers, users have asked for a "show me where the paper says this" button because the animation is so compelling they trust it without checking sources.
Author's perspective
The author agrees with Karpathy's view. As LLM capabilities grow, content production and presentation become increasingly flexible: the same idea can be rendered as text, video, serious tone, or humor, letting each consumer choose their preferred format. The author has been experimenting in this direction with a news agent called Wink Pings, which filters and reorganizes messy information into personalized, easy‑to‑understand styles. Currently text‑focused, the product will expand to images and video as inference costs continue to fall.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
AI Engineering
Focused on cutting‑edge product and technology information and practical experience sharing in the AI field (large models, MLOps/LLMOps, AI application development, AI infrastructure).
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
