Development·EASYHUB JOURNAL
DeepSeek open-sources Ascend inference components and sparse-attention kernels in FlashMLA

What changed
DeepSeek documented and open-sourced inference infrastructure and core components for Huawei Ascend on September 30, including FlashMLA prefill and decoding kernels for DeepSeek Sparse Attention. The project reports 410 and 360 TFLOPS under typical DeepSeek V4.1 workloads, corresponding to about 95% and 83% of theoretical peak. Those performance figures are DeepSeek's own measurements.
- Original title
- A Deep Dive Into the Ascend Sparse Attention Forward Kernel
- Source
- DeepSeek · github.com
- Topic
- Development
- Source month
- 2026-09
This is a concise EasyHub summary of the linked source, not the full report or original reporting. Availability and preview conditions are described in the summary and original.
Summary page published · Updated · Editorial information