Zero-Bug Learning System in 3 AI Dialogues: Seed-2.1-Pro's Multimodal Coding Demo
The author built an open-source knowledge learning system called Knowledge Dance using doubao-seed-2.1-pro-0915 via Trae IDE in three dialogue iterations over three hours, featuring automated research, PPT generation, quiz games, review reports, and knowledge base integration, showcasing the model's multimodal coding and self-debugging via screenshot analysis.
Motivation: Learning Anxiety in the AI Era
The author describes the challenge of keeping up with rapidly evolving knowledge in the LLM era. Traditional "find-read-note" workflows are too slow, while fragmented consumption lacks structure. As a developer, the author wanted an intelligent system that could automatically research topics, plan learning paths, generate structured courseware (PPTs), provide immediate quiz-based feedback, produce review reports, and extend learning with latest web resources — all while persisting progress across sessions.
Five-Step Learning Loop Design
Before prompting the model, the author defined a five-step learning cycle:
Path Planning: Extract key points from user input, search the web, and generate a multi-day learning plan with goals, core concepts, and focus areas.
Content Delivery: Present material as clear, page-by-page PPTs rather than dense text.
Instant Quiz: After each PPT, generate multiple-choice questions with immediate correct/incorrect feedback, explanations, and linked knowledge points — gamified like level progression.
Review Report: Analyze quiz results to score mastery, list weak points, summarize knowledge, and suggest next steps.
Horizon Expansion: For weak areas, search the web for latest tutorials, tools, and frontier developments.
Two long-term features were added: a floating LLM chat that streams explanations grounded in the day's courseware, and a multi-knowledge-base system supporting PDF/Word/Markdown/TXT uploads, vector storage, and selective retrieval for PPT/quiz generation.
Development Stack and Tooling
The project uses
Python FastAPI + LangChain DeepAgents + doubao-seed-2.1-pro-0915(via Volcano Engine Ark API) for the backend, and later Vite + React for the frontend after a separation refactor. The IDE is ByteDance's Trae with the seed model integrated.
Three-Round Iteration Process
Round 1: Core Loop Implementation
The initial prompt defined the agent role, tech stack, eight core capabilities (research→plan→PPT→quiz→review→expand→knowledge-base→persistence), and provided the API key and base URL. The model produced a working server-rendered version covering the full loop.
Round 2: Persistence and Multi-Knowledge-Base
The author requested: (1) save generated learning content to history so switching topics doesn't lose progress, with resume capability; (2) allow creating multiple named knowledge bases, each with its own documents, and let users select one or more bases when generating content. The model implemented both, also handling unstated details like restart recovery and backward compatibility with old data.
Round 3: Frontend-Backend Separation and Visual Upgrade
The author asked to refactor to a decoupled FastAPI + Vite/React architecture and upgrade the UI style referencing a target site (sitor.cc), while preserving all existing functionality. The model performed the structural rewrite and ran full regression tests, confirming zero regressions.
Total elapsed time: ~3.5 hours. The final architecture diagram and screenshots show a polished, feature-complete application.
Multimodal Coding and Self-Debugging via Screenshots
The author highlights Seed-2.1-pro-0915's multimodal coding loop: after writing code, the model automatically starts frontend and backend services, opens a real browser, renders pages, takes screenshots, and visually inspects them like a human frontend engineer + QA. This caught issues pure text debugging misses:
Semantic button placement: "Download PPT" button appeared on the plan page before any PPT existed. The model moved it to the player page after generation.
Responsive layout breakage: Narrow viewport caused knowledge-base two-column layout to collapse; the model added media queries to stack columns vertically.
Visual polish: When mimicking sitor.cc's aesthetic, the model compared screenshots to judge whether colors were "restrained and premium," ultimately settling on paper-texture background, ink-tone headings, dark gradient hero, unified brand blue, and a consistent linear SVG icon set replacing emoji.
All fixes were applied autonomously in the write→screenshot→inspect→patch loop, with near-first-pass success.
Live Demo: Learning RSI (Recursive Self-Improvement)
The author tested the system on the trending topic "RSI" (Recursive Self-Improvement) mentioned in recent Claude Opus 5.2 coverage. Steps:
Enter topic, enable deep-research mode, click "Generate Learning Plan".
System shows real-time progress ("web retrieval, summarizing materials...") and after ~10 minutes outputs a detailed multi-day plan.
Day 1 PPT: 14 slides covering RSI definition, background, context; supports keyboard navigation, dot pagination, and local download.
Quiz phase: 5 multiple-choice questions drawn from PPT content; each answer immediately shows full explanation and linked concept.
Review report: mastery score, weak points, summary, next-step suggestions.
Expansion: web search for latest materials on weak concepts.
History persistence: incomplete sessions remain in history for later continuation.
The end-to-end flow delivers "research → course → exam → grading → review → expansion" in minutes — work that would previously take a small team a full day.
Open Source and Reflection
The project is open-sourced at https://github.com/TangBaron/knowledge-dance. The author's key takeaway: the barrier from "idea" to "working product" is rapidly disappearing. However, the meta-skill of "what to learn, how to learn, and how to verify mastery" becomes more valuable precisely because structured knowledge acquisition is now cheap. The tool paves the road and supplies feedback; the learner must still walk the path.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Fun with Large Models
Master's graduate from Beijing Institute of Technology, published four top‑journal papers, previously worked as a developer at ByteDance and Alibaba. Currently researching large models at a major state‑owned enterprise. Committed to sharing concise, practical AI large‑model development experience, believing that AI large models will become as essential as PCs in the future. Let's start experimenting now!
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
