Karpathy: Replace Raw AI Text with Diagrams, Interactive Pages & Videos

Andrej Karpathy argues that AI now generates output faster than humans can comprehend, proposing four formats — simplified technical English, diagrams, interactive web pages, and explanatory videos — to make AI outputs verifiable, while warning that richer formats make quality control harder.

AI Engineering
AI Engineering
AI Engineering
Karpathy: Replace Raw AI Text with Diagrams, Interactive Pages & Videos

Andrej Karpathy recently posted a thread highlighting a growing mismatch: models can produce tens of thousands of lines of code, research reports, or complex designs in minutes, but humans still need to judge what the model did, whether the logic holds, and whether the result is usable. He proposes four ways to transform model output into forms that are easier to understand and verify.

1. Constrain text with ASD‑STE100

Standard models tend to use long sentences, abstract vocabulary, and vague phrasing. ASD‑STE100 (Simplified Technical English), originally developed for aircraft maintenance manuals, restricts vocabulary, sentence structure, and expression so that each sentence is clear, direct, and unambiguous. A prompt such as "Explain this in 80% ASD‑STE100 style" reduces linguistic noise and makes technical documents easier to audit.

Example of ASD‑STE100 constrained output
Example of ASD‑STE100 constrained output

2. Generate diagrams

Complex relationships — system architecture, call chains, causal links — are hard to convey in linear text. Asking the model to produce flowcharts, architecture diagrams, or concept maps lets readers see how nodes connect, how data flows, and where problems may arise. This is essentially what UML aimed to do, but previously required manual drawing; now an LLM can generate a diagram directly from the specific content at hand.

Generated architecture diagram example
Generated architecture diagram example

3. Generate interactive web pages

When a problem involves state, parameters, or dynamic changes, static text and images are insufficient. The model can output HTML with buttons, sliders, and animations that show how results change when parameters are adjusted or how a system behaves in different states. This turns "reading a conclusion" into "hands‑on verification." A community member noted that interactive pages let users tweak parameters and immediately see causal relationships; another suggested a "guess‑then‑see" mode where the result is hidden until the user predicts the outcome.

Interactive web page with sliders and buttons
Interactive web page with sliders and buttons

4. Generate explanatory videos

Algorithms, physical processes, and abstract concepts all have a temporal dimension. Video can walk through the steps with animation, narration, and highlights. Karpathy suggests asking the model to create 3Blue1Brown‑style explanatory videos. One user generated a 4‑minute video on recursive self‑improvement using Opus 5.5; the source HTML was 13.8 MB but rendered to 34 GB. Karpathy endorsed it. Another user produced a stochastic‑calculus video in a similar style.

3Blue1Brown‑style video frame on recursive self‑improvement
3Blue1Brown‑style video frame on recursive self‑improvement

The underlying shift: disposable software

These approaches share a common driver: the cost of "disposable software" — custom web pages, animations, or videos built for a single question — is approaching zero. Previously, creating a bespoke visualization for one user was uneconomical; now a model can generate it on demand, tailored to the user's knowledge level, and discard it after use. The takeaway: stop forcing yourself to read every word of raw model output; let the model re‑express it in a format that lowers the comprehension cost.

Warning: richer formats make quality control harder

Easier‑to‑consume formats also make errors harder to spot. A flawed paragraph can be scanned quickly; a polished video with animation, narration, and visual effects distributes attention, so mistakes hide inside attractive packaging. Commenters observed that a good‑looking erroneous video can be more persuasive than a flawed text. In interactive papers, users have asked for a "show me where the paper says this" button because the animation is so compelling they trust it without checking sources.

Author's perspective

The author agrees with Karpathy's view. As LLM capabilities grow, content production and presentation become increasingly flexible: the same idea can be rendered as text, video, serious tone, or humor, letting each consumer choose their preferred format. The author has been experimenting in this direction with a news agent called Wink Pings, which filters and reorganizes messy information into personalized, easy‑to‑understand styles. Currently text‑focused, the product will expand to images and video as inference costs continue to fall.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

quality controlinteractive visualizationhuman-AI interactionKarpathyAI output formatsASD-STE100disposable softwareexplanatory videos
AI Engineering
Written by

AI Engineering

Focused on cutting‑edge product and technology information and practical experience sharing in the AI field (large models, MLOps/LLMOps, AI application development, AI infrastructure).

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.