Tagged articles

heterogeneous inference

2 articles · Page 1 of 1
Architects' Tech Alliance
Architects' Tech Alliance
Sep 17, 2026 · Artificial Intelligence

Moore Threads MTT C256 Supernode: 256-GPU Single-Layer Scale-Up Architecture & S5000 Chip Deep Dive

Moore Threads unveils the MTT C256 supernode at WAIC 2026, packing 256 GPUs into a single Scale-up network layer with compute-switch integration, breaking the 64-GPU limit. The MTT S5000 chip delivers 95,920 tok/s FP8 throughput, native FP8 KV Cache, and CUDA-source-level MUSA compatibility for heterogeneous prefill/decode inference.

FP8MTT C256MTT S5000
0 likes · 6 min read
Moore Threads MTT C256 Supernode: 256-GPU Single-Layer Scale-Up Architecture & S5000 Chip Deep Dive
Machine Heart
Machine Heart
Jul 3, 2026 · Artificial Intelligence

Avoiding Pitfalls in Heterogeneous Token Factories: Industry‑Level Design Practices for Cross‑Hardware LLM Inference

The article analyzes a recent multi‑institution paper that maps the design space of heterogeneous Prefill‑Decode LLM inference, identifies three core boundary decisions, presents nine deployment best practices, and validates them with a production token‑factory case on MuXi C600 and NVIDIA Hopper GPUs.

KV CacheLLMdeployment best practices
0 likes · 11 min read
Avoiding Pitfalls in Heterogeneous Token Factories: Industry‑Level Design Practices for Cross‑Hardware LLM Inference