TL;DR
A published report says Huawei trained Pangu Pro, described as a 505-billion-parameter model, without Nvidia accelerators. The available material provides no technical report, hardware inventory or named supply-chain evidence, leaving the model scale and Nvidia-free account unverified.
The original report has linked Huawei Pangu Pro to a 505-billion-parameter training run conducted without Nvidia accelerators, while its headline suggests that conflicting supply-chain evidence may challenge that description. No technical report, chip inventory, supplier records or independent audit accompanied the available material, so the central claims remain unverified.
The reported development contains two main assertions. One describes Pangu Pro as a model with 505 billion parameters that was trained without Nvidia hardware. The other indicates that unspecified supply-chain information tells a different or more complicated story. The available material does not identify the records, suppliers or components behind that qualification.
The parameter claim also lacks details needed to interpret the model’s scale. It is unknown whether 505 billion refers to a dense model, the total capacity of a mixture-of-experts system, or the number of parameters active during each operation. Those configurations can require very different amounts of computing power, memory and communication capacity.
No accelerator model, cluster size, training duration, data volume, architecture, computing budget or evaluation results were provided. The phrase “without Nvidia” is similarly undefined: it could describe only the main training run, or it could claim that Nvidia equipment was absent across experiments, training, evaluation and deployment. Without a documented hardware inventory, readers cannot determine which interpretation was intended.
Huawei Pangu Pro: 505 billion parameters without Nvidia?
A published account describes a frontier-scale, Nvidia-free training run. Yet no technical report, hardware inventory, supplier record or independent audit accompanies the available material. The headline is striking; the evidence remains incomplete.
“Reported” is not the same as technically established.
Three statements, three evidence gaps
The available account combines a model-size claim, a hardware claim and a supply-chain qualification. Each requires different documentation, and none can be confirmed from a headline alone.
“505 billion parameters”
The figure could describe a dense model, total mixture-of-experts capacity or parameters active per operation. Those interpretations imply very different compute and memory needs.
✗ Architecture absent“Without Nvidia”
The scope is undefined. It may refer only to the main run—or claim Nvidia hardware was absent from experiments, evaluation, deployment and supporting infrastructure.
~ Scope undefined“Tells a different story”
No supplier, record or component is named. The alleged complication could concern fabrication, memory, packaging, networking, software or earlier equipment.
✗ Evidence unnamedOne number, radically different systems
A dense 505B model may activate nearly all parameters for each token. A mixture-of-experts model can store hundreds of billions while activating only a fraction. Without an architecture table, direct comparisons are unreliable.
A model is trained by a stack, not a logo
A frontier-scale run depends on synchronized accelerators, high-bandwidth memory, advanced packaging, networking, software and reliable facilities. Processor branding alone cannot establish technological independence.
Why it would matter
If verified, the run would be evidence that a Chinese technology company can execute a frontier-scale workload without relying on Nvidia training accelerators—significant amid export controls and constrained chip access.
Why the caveat matters
A domestically branded accelerator may still rely on foreign-linked intellectual property or components elsewhere in the stack. The available material does not reveal which layer, if any, prompted the supply-chain qualification.
What is present—and what is missing
Verification requires evidence matched to each assertion. The current material offers attributed claims, but not the technical and documentary trail needed to reproduce or independently audit them.
| Evidence item | Why it matters | Available material | Verification value |
|---|---|---|---|
| Model card | Confirms architecture, scale and evaluation scope | ✗Not disclosed | Essential |
| Total + active parameters | Distinguishes dense capacity from MoE routing | ✗Not separated | Essential |
| Accelerator inventory | Identifies chip models, quantity and cluster topology | ✗Not supplied | Essential |
| Training logs | Documents duration, utilization, failures and compute | ✗Not supplied | High |
| Supplier records | Tests the claimed supply-chain contradiction | ✗No source named | Essential |
| Headline attribution | Establishes that the claim was publicly reported | ✓Present | Limited |
| Independent audit | Connects architecture, hardware and suppliers | ✗None cited | Strongest |
✓ Present ~ Partial or ambiguous ✗ Not disclosed in the available material
How the dispute could be settled
The shortest route from headline to credible conclusion is a traceable chain connecting the model architecture to the training run, the physical cluster and independently checkable supply-chain records.
What readers should ask next
The core issue is not whether the report is possible. It is whether the available evidence is sufficient to distinguish a verified engineering achievement from an attributed but underspecified claim.
Did Huawei confirm 505 billion parameters?
No Huawei technical document, model card or architecture disclosure confirming that count appears in the available material.
Was training definitely Nvidia-free?
No definitive proof was supplied. There is no accelerator list, cluster inventory or independent audit identifying the training hardware.
What does 505 billion mean?
It may describe total stored parameters or parameters active during each operation—an important distinction for mixture-of-experts systems.
What contradicts the supply-chain account?
The material names no supplier, record or component and does not establish that Nvidia components were present.
China’s AI Hardware Independence Test
If verified, the reported training run would offer evidence that a Chinese technology company can operate a frontier-scale AI workload without relying on Nvidia training accelerators. That would matter as export controls and chip availability shape the computing resources available to Chinese AI developers.
The claim also concerns more than the processor brand. Large training clusters depend on high-bandwidth memory, advanced packaging, networking, fabrication tools, software, power and cooling. A domestically branded accelerator could still rely on foreign-linked components or intellectual property elsewhere in that stack. The suggested supply-chain conflict may concern any of those layers, but the source material does not identify one.

HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
- Architecture: NVIDIA Volta GV100 with CUDA and Tensor Cores
- Memory: 32GB HBM2 ECC with 900 GB/s bandwidth
- Interface: PCIe 3.0 x16 with 250W TDP
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Stack Behind Pangu Pro
Training a model at the reported scale would require a coordinated system of accelerators, memory, networking and distributed-training software. Processor quantity alone would not establish capability: chip yields, interconnect performance, compiler maturity and cluster reliability can all affect whether a large run finishes within a practical period and budget.
The distinction between total and active parameters is also central. A mixture-of-experts model may contain hundreds of billions of total parameters while activating only a fraction for each token. Without an architecture description, the 505-billion figure cannot be used to calculate the run’s likely computing needs or compare Pangu Pro fairly with other systems.

Accelerate Everything with Tensor Cores: A Developer’s Guide to High-Performance AI, Efficient Training, and Scalable Models
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Missing Proof Behind Both Claims
Neither the reported model scale nor the hardware account has been independently established in the material provided. There is no disclosed model card, architecture table, training log, cluster inventory, supplier document or third-party audit connecting Pangu Pro to the claimed run.
It is also unclear what the alleged supply-chain discrepancy concerns. Possible areas include accelerator fabrication, memory, packaging, networking equipment, software dependencies or hardware used during earlier experiments. These are possible interpretations, not confirmed findings. The material does not establish that Nvidia components were present, only that the headline questions or qualifies an Nvidia-free characterization.
large scale neural network training servers
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Documents Needed to Verify Pangu
Verification would require Huawei or the report’s publisher to disclose the model architecture, total and active parameter counts, accelerator type, cluster size, training methodology and benchmark results. A clear definition of “without Nvidia” would also need to state whether it covers preliminary experiments, the main run, evaluation and deployment.
The supply-chain side would require named components, suppliers or records that can be checked independently. Until such material appears, the responsible reading is that a large Nvidia-free training run has been reported, while both its technical basis and the claimed supply-chain contradiction remain open.

HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
- Architecture: NVIDIA Volta GV100 with CUDA and Tensor Cores
- Memory: 32GB HBM2 ECC with 900 GB/s bandwidth
- Interface: PCIe 3.0 x16 with 250W TDP
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
Did Huawei confirm that Pangu Pro has 505 billion parameters?
The available source material attributes the figure to a report but does not include a Huawei technical document, model card or architecture disclosure confirming the 505-billion-parameter count.
Was Pangu Pro definitely trained without Nvidia chips?
No definitive proof was supplied. The report’s headline makes the Nvidia-free claim, but no accelerator list, cluster inventory or independent audit identifies the hardware used for training.
What does 505 billion parameters mean?
It describes a claimed measure of model capacity, but its meaning depends on the architecture. The figure could represent all stored parameters or parameters active during each operation. That distinction is especially relevant for a mixture-of-experts model.
What supply-chain evidence contradicts the account?
The available material does not say. It names no supplier, record or component, and does not establish whether the issue concerns processors, fabrication, memory, packaging, networking, software or earlier equipment.
What evidence could settle the dispute?
A detailed technical report and hardware inventory, supported by training logs and independently checkable supplier records, could clarify the model’s scale, the accelerators used and the scope of the Nvidia-free description.
Source: Thorsten Meyer AI