
What happened
DeepSeek open-sourced six tools for Huawei's Ascend chips, including TileLang, DeepGEMM Ascend and DeepEP Ascend, with Huawei backing the work. It also optimized systems carrying 128 Ascend 950 chips.
Why it matters
Developers get a simpler alternative to Nvidia's CUDA for writing AI code, which may lower the cost of building on Huawei hardware.
What to watch
Whether the toolset performs far beyond Nvidia GPUs in practice — DeepSeek reports FlashMLA hit 95% of theoretical performance on Ascend 950 and DeepSelect ran 2x to 20x faster than torch.topk.
WHO IT HITSAI developers and infrastructure engineers who write performance-critical code for AI chips gain a Huawei-compatible path that avoids CUDA lock-in, while data-center teams weighing Ascend hardware may see more mature software support.
Summaries like this, in your inbox every morning.
The release arrived with a concrete technical marker: TileLang added Huawei Ascend 950 as an official backend on September 30, 2026, joining Nvidia CUDA, AMD ROCm and Apple Metal. Beyond the language itself, DeepSeek shipped libraries that cover matrix multiplication, inter-NPU communication for Mixture of Experts models, and attention processing used in DeepSeek-V4.1 inference, plus a set of dozens of high-speed kernels that automatically detect whether an Nvidia GPU or Huawei NPU is running.
Huawei fully supported the development effort, and the two companies also worked on optimizing a "supernode" configuration carrying 128 Ascend 950 chips, tuning both compute and chip-to-chip communication. DeepSeek's stated goal is a high-level language that is general-purpose, easy to program and still able to extract hardware performance for AI accelerators.
The competitive read is that this is one of the clearest efforts yet to build a software layer that lets AI workloads run on Huawei silicon without rewriting code written for Nvidia's ecosystem, though the practical test will be whether real-world workloads match the reported numbers — FlashMLA reaching 95% of theoretical performance on some operations and DeepSelect's 2x to 20x speedup over torch.topk. For teams currently tied to CUDA, the toolchain may offer a lower-friction migration path, but only if the performance claims hold outside DeepSeek's own benchmarks.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
DIGITIMES reported the SiC substrate market is finally seeing signs of recovery, after prices collapsed when w…

HENNGE said on October 1 it set up HENNGE AI, a subsidiary with only two directors and no other employees, whe…

Pershing Square sold its entire Alphabet stake and topped off Microsoft while it sat about 20% below its Q2 hi…

Aolani will integrate Karman's power orchestration platform, built on a custom NVIDIA Jetson Orin Nano and tar…

CloudNC raised $20 million led by Nimble Ventures, with Calculus Venture Capital, Entrepreneur First and LM Ve…

Solomon Asamoah laid out a framework for Ghana's AI infrastructure, arguing data-centre approvals should start…
