Moore Threads MTT C256 Supernode: 256-GPU Single-Layer Scale-Up Architecture & S5000 Chip Deep Dive

Moore Threads unveils the MTT C256 supernode at WAIC 2026, packing 256 GPUs into a single Scale-up network layer with compute-switch integration, breaking the 64-GPU limit. The MTT S5000 chip delivers 95,920 tok/s FP8 throughput, native FP8 KV Cache, and CUDA-source-level MUSA compatibility for heterogeneous prefill/decode inference.

Architects' Tech Alliance
Architects' Tech Alliance
Architects' Tech Alliance
Moore Threads MTT C256 Supernode: 256-GPU Single-Layer Scale-Up Architecture & S5000 Chip Deep Dive

At WAIC 2026, Moore Threads publicly demonstrated the MTT C256 supernode for the first time, turning its roadmap into a shipping product. The system organizes 256 GPUs (MTT S4000/S3000 series) within a single Scale-up network fabric, using a compute-switch integrated high-density design that eliminates the bandwidth loss and forwarding latency inherent in traditional multi-layer network topologies. It targets three core AI workloads: model training as the building block for 100,000-GPU clusters supporting pre-training, post-training, and reinforcement learning; inference service with native large-scale MoE adaptation and million-token context support; and inference decoding for high-concurrency, real-time scenarios such as AI coding assistants.

Single-Layer 256-GPU Scale-Up Breakthrough

The central innovation is a single-layer Scale-up network that lifts the industry-standard 64-GPU limit directly to 256 GPUs. This fundamentally mitigates the bandwidth attenuation and communication latency problems of multi-layer architectures. A single cabinet achieves 128-GPU full mesh; two standard cabinets complete the 256-GPU single-layer interconnect.

Moore Threads Full-Family Chip Portfolio

The product line centers on full-function GPUs. The MTT S5000 is in volume production, while the next-generation "Huagang" architecture is under planning. The "Changjiang" SoC adopts a unified memory architecture, enabling zero-copy data movement between CPU and GPU across graphics, AI, and other workloads to maximize task execution speed.

MUSA Unified System Architecture

MUSA (Moore Threads Unified System Architecture) is a self-developed fused GPU hardware/software stack covering unified chip architecture, instruction set, programming model, software runtime libraries, and driver framework, aiming to deliver high-performance parallel computing across diverse scenarios.

MTT S5000 Technical Analysis

The MTT S5000 serves as the core hardware for Prefill-as-a-Service compute pools in heterogeneous PD (prefill/decode) separation architectures, pairing with GPU A (decode generation pool). Its key technical attributes:

High-density FP8 compute : Matches prefill's compute-intensive profile; 8-card single-node throughput reaches 95,920 tok/s at 64K input length.

Long-context stability : Across 64K–400K input range, P50 TTFT (time-to-first-token) remains stably within a 6-second SLO.

Native FP8 KV Cache : Precision-aligned with the decode pool, supports page-granularity direct transfer, eliminating format-conversion overhead.

MUSA ecosystem CUDA compatibility : Source-level migration; supports mainstream frameworks including PyTorch, SGLang, and vLLM.

GPUDirect RDMA + Mooncake : 400G RDMA networking enables efficient cross-node KV Cache transfer.

Full-precision support : FP8 through FP64; among the earliest domestic GPUs to support FP8.

The article positions the MTT S5000 as a general-purpose AI accelerator card purpose-built for the Prefill compute pool in disaggregated inference architectures.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

FP8supernodeMUSAMoore Threadsheterogeneous inferenceMTT C256MTT S5000Scale-up network
Architects' Tech Alliance
Written by

Architects' Tech Alliance

Sharing project experiences, insights into cutting-edge architectures, focusing on cloud computing, microservices, big data, hyper-convergence, storage, data protection, artificial intelligence, industry practices and solutions.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.