Can We Watch a 535B Model Train Live? Opening the Black Box of LLM Training

Stanford professor Percy Liang has launched the public, three‑month training of the 535B‑parameter Marin model, detailing its 18.75 trillion‑token dataset, 792 GB200 GPUs, 2.7e24 FLOPs compute, scaling‑ladder pre‑runs, and live monitoring links, sparking worldwide community interest.

Machine Heart
Machine Heart
Machine Heart
Can We Watch a 535B Model Train Live? Opening the Black Box of LLM Training

Stanford professor and Simile AI founder Percy Liang announced that the Marin laboratory will publicly train the 535B‑A23B large language model, making the entire training process observable in real time.

The model contains 5.35 trillion total parameters and 230 billion activation parameters. The team prepared 18.75 trillion tokens of training data and deployed eleven GB200 NVL72 systems, equivalent to roughly 792 GB200 GPUs. Training is expected to run for about three months, consuming approximately 2.7 × 10^24 FLOPs, followed by a post‑training phase.

Before the hero run, Liang’s team executed a four‑stage "Scaling Ladder" ranging from a 1.6 B‑parameter model (48 B tokens) up to a 27.7 B‑parameter model (926 B tokens). The ladder serves two purposes: to uncover and debug potential issues early and to predict the performance of the full‑scale run.

All training metrics—including data composition, loss curves, model states, and intermediate predictions—are streamed publicly. View the data composition at https://storage.googleapis.com/marin-public/held/harrier-k40-cluster-overview/2026.08.18/index.html?revision=uniform-sampling, monitor the live run on Weights & Biases at https://wandb.ai/marin-community/marin_moe/reports/535B-A23B-18T-Token-Hero-Run-Scaling-Ladder--VmlldzoxNzc2MDM5Ng, and explore detailed engineering information on GitHub at https://github.com/marin-community/marin/issues/8435.

The community has reacted enthusiastically, describing the effort as true open‑source spirit and likening it to opening a previously sealed "dark room" of AI labs. One observer noted an unexpected norm spike around step 500, which was later identified as caused by the router_bias term in the mixture‑of‑experts (MoE) architecture rather than a gradient‑trained parameter.

Liang also teaches the course "CS336: Language Models From Scratch," aimed at teaching students how to build large language models from the ground up; the lecture series is available at https://www.youtube.com/playlist?list=PLoROMvodv4rMqXOcazWaTUHhq-yembLCV.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

large language modelmodel trainingopen researchGPU computeMarinscaling ladder
Machine Heart
Written by

Machine Heart

Professional AI media and industry service platform

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.