NEWS / SEP.2026
DeepSeek makes it easier to port AI computations to Huawei chips
On September 30, 2026, DeepSeek released six software components adapted to Huawei Ascend chips, with support for Ascend 950. Some preserved interfaces give teams a foundation for porting, whose benefits in production must be tested on a complete model.

DeepSeek makes it easier to port AI computations to Huawei chips
On September 30, 2026, DeepSeek made public six software components adapted to Huawei Ascend, covering compilation, computations and communications between accelerators. The South China Morning Post reports this release and links it to the announcement on DeepSeek’s official WeChat account.
The announcement brings together TileLang, DeepGEMM-Ascend, DeepEP-Ascend, TileKernels, FlashMLA and DeepSelect. These components adapt earlier tools to a different hardware architecture. TileLang is a community project hosted by tile-ai. Its guide credits DeepSeek members primarily with the initial development of Ascend 950 support, with integration by the TileLang community and collaboration from Huawei.
Official native Ascend 950 support in TileLang is dated September 30, 2026. The project’s changelog already mentioned Huawei Ascend adapters on September 29, 2025, in an external project. The new support therefore builds on work already underway on Ascend.
Reusing software calls on Ascend
TileLang allows computations to be written in Python-like syntax. Its compiler transforms this description into programs adapted to the chip. For Ascend 950, the project announces native code generation and automated scheduling and synchronization. These functions organize the execution of computations and coordinate operations.
DeepSeek says that DeepGEMM-Ascend preserves DeepGEMM’s interfaces for matrix multiplication, a core operation in AI computing. A team can reuse some of its software calls and development practices. The library supports BF16, FP8 and FP4, numerical formats with different levels of precision.
TileKernels documents the same Python interfaces on NVIDIA and Huawei, with automatic selection of the implementation matching the hardware. The library covers, among other things, routing data to a model’s experts, quantization, which changes their numerical precision, and data transformations.
FlashMLA provides sparse attention computations for Ascend 950, from initial context processing to the successive generation of tokens, the units of text handled by the model. This attention focuses computation on a selection of elements. DeepSelect performs TopK selection, retaining the K highest-ranked elements, particularly for sparse attention and sampling.
DeepEP-Ascend handles exchanges between accelerators for distributed training and inference, the latter referring to using the model to produce responses. It covers, among other things, sending and recombining data for models with experts. These components cover part of the infrastructure needed to run a model.
For developers, laboratories and AI operators, reusing these interfaces could reduce the rework needed to use Huawei and broaden the choice of hardware, provided the tools are adopted and maintained. Public-sector buyers could compare porting effort, reproducibility and maintenance when evaluating infrastructure. The amount of code that can actually be reused will depend on each application.
Adaptations to verify before deployment
DeepGEMM-Ascend is developed and validated on the Ascend 950 series. Its documentation requires, among other things, the Huawei CANN 9.20 toolkit, torch_npu and Python 3.10 or later. TileKernels requires Ascend 950 and CANN 9.2.0 or later. These prerequisites maintain a dependence on Huawei equipment and tools.
The shared interfaces come with technical differences. DeepGEMM-Ascend documents scaling factors and alignment constraints specific to Ascend. Some FlashMLA operations are still exclusive to CUDA, NVIDIA’s software environment. DeepSelect accepts bfloat16 on Ascend, while float32 is supported only by CUDA.
The published measurements for DeepEP-Ascend concern Ascend 950DT with CANN 9.2.0 and a proof-of-concept version of the HDK software kit, provided to DeepSeek and configured manually. This configuration is not publicly distributed. The repository states that the commercial kit recommended for Atlas 850E is expected around October 15, 2026, subject to release by Huawei.
The published results concern specific computations or communications. They do not allow the overall savings from a complete training run at comparable precision to be quantified.
Validation in production would require reproducible tests by external teams on a complete model, with comparable and documented hardware, software versions and precision. They should measure throughput, result accuracy, stability and energy consumption, as well as porting time, maintenance and total cost. The public code now provides a foundation for starting these tests.