DeepSeek Releases 6 Ascend AI Kernels: TileLang, DeepGEMM, FlashMLA & More
DeepSeek open-sources six low-level libraries for Huawei Ascend 950, covering operator DSL, matrix multiplication, MoE communication, sparse attention, and TopK selection, with reported near-peak hardware performance but requiring version-specific integration.
On September 30, DeepSeek-affiliated teams released or updated a batch of software components targeting Huawei Ascend 950 chips. The six projects are not new models or ready-to-use applications, but low-level tools spanning operator development, matrix computation, MoE communication, attention, and TopK selection. Together they form a complete software stack that mirrors DeepSeek's previously open-sourced NVIDIA-oriented components, enabling model teams to build and optimize training and inference pipelines on Ascend hardware.
Six Projects at a Glance
TileLang – A domain-specific language (DSL) for writing high-performance operators with Pythonic syntax, built on TVM. The update adds an official Ascend 950 backend. The repository belongs to the tile-ai project, not solely maintained by DeepSeek.
TileKernels – A library of operators implemented in TileLang, covering MoE routing, quantization (FP8/FP4 with fused SwiGLU), Engram gating (fused RMSNorm), manifold hyper-connection (mHC) with Sinkhorn normalization, RoPE, and random number generation. Acts as reusable operator building blocks.
DeepGEMM-Ascend – Handles compute-intensive matrix multiplication, supporting BF16, FP8, and FP4 formats. On Ascend 950 DT, dense GEMM reaches 99.8% of theoretical BF16 peak, 99.5% for FP8, and 98.3% for FP4 (official README figures, specific test conditions). Provides MQA logits and MegaMoE operators. Abstracts Ascend's MAD primitive, hiding fractal matrix layout, alignment constraints, and address calculation, while adding sparse data loading and coroutine-based pipelining.
DeepEP-Ascend – Expert-parallel communication library for MoE models, managing token dispatch and combine all-to-all operations. Aligns with the public Buffer API of the NVIDIA version. Includes EPBuffer (FP8 dispatch, delayed epilogue), with planned PPBuffer for pipeline parallelism, BucketBuffer for context/data parallelism, and Engram remote memory access. Kernels run on HCCL/HCOMM, UBMEM, and URMA primitives, with runtime compilation via DeepJIT. On Ascend 950 DT super-node external Clos network, dispatch sustained bandwidth hits 90–95% of physical payload limit for EP scale ≤32; combine still under optimization.
FlashMLA – High-performance attention operator library supporting DeepSeek-V4.1 inference on both NVIDIA GPU and Ascend NPU. Core is token-level sparse attention (DeepSeek Sparse Attention, DSA): Lightning Indexer selects top-k tokens per query, attention computed only on selected tokens, drastically reducing long-context compute. Ascend release includes sparse prefill and decoding operators with FP8/FP4 KV cache (V4.1: 528 bytes/token, FP4: 288 bytes/token). On Ascend 950, prefill peaks at 410 TFLOPS (95% theoretical peak), decoding at 360 TFLOPS (83% peak). Accompanied by a technical report detailing algorithms and optimizations.
DeepSelect – Focused high-performance TopK operator. Targets two hot paths: Lightning Indexer token selection in DSA (top-k=512 in DeepSeek V4) and sampler TopK sampling from ~128K vocabulary during generation. Both execute per generated token.
TileLang: The Foundational DSL
TileLang is the keystone of this release. It provides a concise DSL for writing high-performance operators in Pythonic syntax, with a compiler backend built on TVM. The Ascend version wraps Ascend C low-level instructions, allowing high-level programming without sacrificing hardware performance. TileLang's backend list now includes NVIDIA CUDA (SM70 to SM120), AMD ROCm, Apple Metal, and Ascend 950. DeepSeek aims to make TileLang a reference for building usable software ecosystems on diverse AI chips.
DeepGEMM-Ascend: Matrix Multiplication at Hardware Limits
Matrix multiplication dominates large-model training and inference. DeepGEMM-Ascend is a port of DeepSeek's established NVIDIA GEMM library, achieving full API compatibility—same package name, interfaces, and development flow—for near-seamless switching. It supports BF16, FP8, FP4 GEMM, MQA logits, and MegaMoE operators. Implementation lightly abstracts Ascend's MAD primitive, hiding fractal layout, alignment, and address details, while layering Ascend-specific sparse loads and coroutine pipelining. Reported peaks on Ascend 950 DT: 99.8% (BF16), 99.5% (FP8), 98.3% (FP4) of hardware theoretical maximum.
DeepEP-Ascend: Communication Backbone for 128-Card Super-Nodes
MoE models live or die by communication efficiency. DeepEP-Ascend mirrors the NVIDIA version's public Buffer API. Its Ascend C kernels sit atop HCCL/HCOMM, UBMEM, and URMA, with DeepJIT runtime compilation. Beyond core EPBuffer (FP8 dispatch, delayed epilogue), the roadmap includes PPBuffer for pipeline parallelism, BucketBuffer for context/data parallelism, and Engram remote memory access. On Ascend 950 DT super-node external Clos network, dispatch sustained bandwidth reaches 90–95% of physical payload ceiling for EP scale ≤32; combine optimization ongoing.
TileKernels: Internal Operator Arsenal Made Public
TileKernels is a library of dozens of deeply optimized operators, all written in TileLang, covering the most common operations in large-model training and inference: MoE routing top-k selection and scoring; per-token/per-block/per-channel FP8/FP4 quantization with fused SwiGLU; Engram gating with fused RMSNorm (forward and backward); manifold hyper-connection (mHC) with Sinkhorn normalization; RoPE rotary position encoding; and random number generation.
FlashMLA: Sparse Attention for Long-Context Inference
FlashMLA powers DeepSeek-V4.1 inference on both NVIDIA GPU and Ascend NPU. Its core is token-level sparse attention (DSA): Lightning Indexer picks top-k tokens per query, attention computed only on those tokens, slashing long-context compute. The Ascend release provides sparse prefill and decoding operators with FP8/FP4 KV cache (V4.1: 528 bytes/token, FP4: 288 bytes/token). Performance numbers are concrete: Ascend 950 prefill up to 410 TFLOPS (95% peak), decoding up to 360 TFLOPS (83% peak). A technical report detailing algorithms and optimization tricks is available.
https://github.com/deepseek-ai/FlashMLA/blob/main/docs/20260930-ascend-prefill-deep-dive-zh.mdDeepSelect: Optimizing the Critical TopK Hot Path
DeepSelect is the smallest project, focusing solely on high-performance TopK. Its two critical use cases are: (1) Lightning Indexer token selection in DSA sparse attention (top-k=512 in DeepSeek V4), and (2) inference sampler TopK sampling from ~128K vocabulary. Both lie on the per-token generation hot path, justifying dedicated optimization.
Caveats and Deployment Guidance
All performance figures (99.8%, 90–95%, 410 TFLOPS) come from project maintainers' public tests or technical reports; the author has not independently verified them on Ascend hardware. Results depend on matrix shapes, expert parallel scale, network topology, data formats, and model versions. Operator microbenchmarks do not directly translate to end-to-end request latency. The repositories are published independently; their dependencies (CANN, PyTorch NPU, Python, compiler toolchains) have not been validated as a unified installation matrix. TileLang and TileKernels both target Ascend 950 but have differing build requirements. Teams with existing Ascend environments evaluating MoE or long-context inference should verify device generation, driver/CANN versions, torch_npu, and compiler requirements, then run end-to-end tests with their own models, input lengths, and request loads. For developers without target hardware or migration tasks, these repos serve as valuable low-level reference implementations rather than immediate deployment solutions.
References
TileLang — github.com/tile-ai/tilelang DeepGEMM-Ascend — github.com/deepseek-ai/DeepGEMM-Ascend DeepEP-Ascend — github.com/deepseek-ai/DeepEP-Ascend TileKernels — github.com/deepseek-ai/TileKernels FlashMLA — github.com/deepseek-ai/FlashMLA DeepSelect —
github.com/deepseek-ai/DeepSelectSigned-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Data STUDIO
Click to receive the "Python Study Handbook"; reply "benefit" in the chat to get it. Data STUDIO focuses on original data science articles, centered on Python, covering machine learning, data analysis, visualization, MySQL and other practical knowledge and project case studies.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
