DeepSeek Trains Models on Huawei Ascend 950 Silicon, Report Says
WHY IT MATTERS
r/LocalLLaMA reported that DeepSeek has now trained models on Huawei's Ascend 950 accelerator. The claim originated as a community post and has not been independently confirmed via official channels.
What Happened
A post on r/LocalLLaMA claims that DeepSeek has trained models on Huawei's Ascend 950 accelerator. The report describes an internal training workload executed on Ascend silicon rather than NVIDIA hardware. The claim has not been confirmed by DeepSeek, Huawei, or any independent benchmarking source. Ascend 950 is a newer generation part than the widely deployed 910B, and public documentation on its training throughput remains sparse.
Why It Matters
If accurate, this is the first credible signal of a frontier-tier Chinese lab running meaningful pretraining on domestic, non-NVIDIA silicon. That would mean export controls are no longer the binding constraint on at least one major lab's training pipeline, changing the cost and supply assumptions that Western operators currently price into frontier model competition. For Chinese labs, Ascend removes single-point dependency on smuggled or stockpiled H100/H800 inventory and gives them a domestic supply chain that scales with local fab capacity rather than US licensing. For Western labs and cloud providers, it erodes the assumption that compute scarcity will keep Chinese frontier capabilities one hardware generation behind. It also reframes Huawei from a networking and inference vendor into a direct competitor to NVIDIA's training stack in the largest addressable AI market outside the US.
Technical Details
Ascend 950 is the successor to the 910B and 910C, using Huawei's Da Vinci architecture with a redesigned interconnect intended to close the scale-out gap that made earlier Ascend parts inefficient for large training runs. Reported specs for the 950 line include higher FP8 and FP16 throughput per die and a proprietary interconnect (HCCS) positioned against NVLink, with the practical ceiling depending heavily on whether Huawei's CANN software stack and collective communication libraries can sustain utilization across thousands of dies. CANN has historically been the weaker layer versus CUDA: kernel coverage, graph compilation, and framework integration (PyTorch, Megatron-style parallelism) all require more operator effort and yield lower MFU on equivalent hardware. The r/LocalLLaMA claim does not specify model size, sequence length, token count, cluster scale, or achieved MFU, which are the numbers that would separate a genuine frontier pretraining run from a fine-tune or limited-scale experiment.
Operational Impact
For builders, the immediate read is not that Ascend is a viable substitute for H100 clusters but that the assumption of a single dominant training stack is weakening. Teams evaluating multi-vendor portability now have a concrete case to point at when justifying abstraction layers — Megatron forks, Triton ports, and framework-agnostic parallelism schemes become more defensible budget line items. If DeepSeek publishes benchmarks, expect inference-focused operators in regulated markets (China, parts of the EU, sovereign AI programs) to start trialing Ascend as a cheaper per-token alternative to NVIDIA for serving, even if training stays on NVIDIA. For anyone running capacity planning, the model to track is not "can Ascend beat H100" but "what is the effective cost per useful FLOP including engineering overhead," which is where CANN historically loses. Workflows that hard-code CUDA-specific kernels or rely on NCCL tuning will need equivalents if teams want to hedge.
SOURCE
SHARE
MORE FROM STUFFINSIDER
Moderna Jumps 110% on Positive Phase 3 Cancer Vaccine Results
Sep 25INDUSTRYAnthropic financial-services Repo Trends on GitHub With 236 Stars
Sep 20INDUSTRYGoogle DeepMind: Gemini Hacked Three Companies in Security Tests
Sep 19INDUSTRYModerna Stock Surges 110% on Positive Phase 3 Cancer Vaccine Results
Sep 15