CHINA AI
INFRASTRUCTURE
Home / Deep dives / ByteDance infrastructure
← Back to Ecosystem Map
COMPANY ECOSYSTEM DEEP DIVE • RESEARCHED 19 SEP 2026

ByteDance is already an AI infrastructure company.

The useful way to model ByteDance is not simply as a consumer internet company that buys GPUs. Its consumer AI demand, Seed research and Volcano Engine cloud business pull it downward into scheduling, heterogeneous accelerators, proprietary networking, storage, custom CPUs, data-centre campuses and grid-scale power infrastructure.

Working classification: Demand-to-infrastructure hyperscaler.
ByteDance starts with enormous application and token demand, then engineers downward through the stack to make that demand economically serviceable. That makes it a strong real-world test of the site's central thesis: heterogeneous compute islands can become more useful when networking, software, scheduling, data systems and physical infrastructure are designed as one system.

The ByteDance machine

AI DEMANDDouyin • Doubao • Seed • Seedance • Coze • Feishu VOLCANO ENGINE / CONTROL PLANEMaaS • cloud • K8s • scheduling • resource pooling HETEROGENEOUS COMPUTEHuawei • Cambricon • IluvatarNvidia / Arm / custom silicon NETWORK FABRICveRoCE • EthLink • RDMAcompute ↔ communication overlap DATA / STORAGEEB-scale HDFS • CloudFS • IcebergPB/day growth • data lake AIDC + POWERWuhu • Wuwei • 220kV substationsSoutheast Asia expansion INTERNAL AI OUTPUTtraining • inference • agents • media EXTERNAL MONETISATIONVolcano Engine • MaaS • enterprise AI
Scroll horizontally to explore the full diagram →
China AI Infrastructure · chinaaiinfra.netlify.app · See accompanying evidence and review date.

What the research changed

1 • It is not merely buying acceleratorsByteDance's Seed infrastructure team explicitly owns distributed training, high-performance inference and heterogeneous-hardware compilation. COMET is deployed in production clusters at tens-of-thousands-of-GPU scale.
2 • Networking is an owned optimisation layerByteDance developed veRoCE for RDMA and EthLink for Ethernet scale-up. Company-reported testing says veRoCE improved LLM training speed 11.2% and All-to-All throughput 48.4%.
3 • Physical infrastructure is now strategicWuhu and Wuwei campuses include dedicated 220kV substations. The first Wuhu phase was described by China Energy Engineering as ByteDance's first East China data-centre campus, with total investment expected around RMB8bn.
4 • CPU is becoming part of the AI stackReuters reports ByteDance is developing Arm and RISC-V custom CPUs as inference and agent workloads increase CPU demand; Arm separately identified ByteDance as a customer of its AI data-centre CPUs.
5 • Domestic GPU diversification is realReuters says Huawei is ByteDance's largest domestic accelerator supplier, followed by Cambricon and Iluvatar CoreX. Iluvatar's 2026 allocation was reportedly doubled to 100,000 GPUs as constraints intensified.
6 • Geography is becoming part of orchestrationInner Mongolia's development commission disclosed talks with ByteDance on compute monitoring, trading and converting regional compute advantages into token output. This is emerging policy/market evidence, not proof of production-scale dynamic scheduling.

The accelerator ecosystem

EDGE
STATUS
WHAT IS ACTUALLY EVIDENCED
WHY IT MATTERS
ByteDance × Huawei
CORROBORATED REPORTING
Reuters sources in September described Huawei as ByteDance's largest domestic AI-chip supplier.
Strongest domestic compute edge.
ByteDance × Cambricon
CORROBORATED REPORTING
Reuters describes Cambricon as the second-largest domestic supplier.
Confirms multi-vendor compute.
ByteDance × Iluvatar CoreX
REPORTED
June talks targeted at least 50k inference chips; September reporting said allocation doubled to 100k units.
Inference demand can pull smaller domestic vendors into hyperscale production.
ByteDance × Kunlunxin
CONSIDERED / NOT CONFIRMED
Reuters reported ByteDance was considering Baidu Kunlunxin in June. No sufficient evidence found here of completed production deployment.
Keep as research gap.
ByteDance × Qualcomm
REPORTED
Reuters, citing Bloomberg, reported an AI-ASIC deal involving millions of chips for data-centre/agent workloads; neither company commented.
Potential custom-silicon route beyond GPUs.
ByteDance × Arm
CONFIRMED BY ARM
Arm CEO identified ByteDance as a customer of its AI data-centre CPUs; Reuters separately reports custom Arm/RISC-V CPU development.
Inference shifts system bottlenecks toward CPU as well as GPU.
HBM is the stress test. The September HBM shortage is not peripheral to ByteDance. Reuters reported price increases across Huawei, Cambricon and Iluvatar and said Iluvatar diverted internally earmarked GPUs to ByteDance. This is precisely where the thesis has to become economic: can ByteDance's scheduler, network, model architecture and heterogeneous sourcing offset enough of the higher component cost?

Networking: ByteDance is building around the chip

veRoCE

ByteDance's high-performance RDMA protocol adds multipath/out-of-order optimisation, SACK retransmission and multipath congestion control. Volcano Engine says IANA assigned it UDP port 4794, helping multi-vendor hardware recognise the protocol.

+11.2%
ByteDance-reported LLM training-speed improvement in testing.
EthLink

ByteDance has published an Ethernet-optimised GPU scale-up interconnect design aimed at low-latency, high-bandwidth communication inside AI racks, complementing scale-out networking.

Scale-up + scale-out
The infrastructure strategy is moving from generic network plumbing toward AI-specific fabric design.

Control plane and useful compute

This is where ByteDance becomes particularly important to our thesis. Its compute-infrastructure roles describe unified scheduling across containers, VMs, online/offline workloads and CPU/GPU resources, including intelligent scheduling across CPU, GPU, memory, network and even power across global data centres. Volcano Engine's public products expose the same direction externally through heterogeneous GPU support, K8s scheduling and unified resource pools.

Interpretation — not a company claim: ByteDance is one of the clearest Chinese examples of heterogeneous islands connected by an increasingly intelligent control plane. It does not need Huawei, Cambricon, Iluvatar, Nvidia, Arm and future custom silicon to be interchangeable. It needs the software/network layer to make each pool usable for the workloads it is good enough to run.

Physical infrastructure and power

Wuhu Yangtze River Delta compute centreChina Energy Engineering says the first phase included four 150MVA split-winding transformers and three 220kV outgoing lines. A 2026 procurement notice shows continued B-phase/7-5 construction activity.
Wuwei compute centreA second ByteDance/Volcano Engine campus in Wuwei has a planned 220kV substation with five 150MVA split-winding transformers and 140 10kV outgoing circuits. Environmental filings were active in 2026.
Thailand / Southeast AsiaThailand approved a $25bn TikTok data-infrastructure expansion in May 2026 across Bangkok, Samut Prakan and Chachoengsao. ByteDance's September financing also supports AI infrastructure including Southeast Asian data centres.

A VNET connection we should keep

Volcano Engine × DYXnet is confirmed. DYXnet—VNET's wholly owned subsidiary—said it signed cooperation with Volcano Engine covering cloud computing, AI, security and network bandwidth. This does not establish that ByteDance is a major VNET colocation customer, and I found no sufficiently strong evidence for that broader claim. But it is a genuine ByteDance/Volcano Engine ↔ VNET ecosystem edge and belongs in the relationship atlas.

Economics: the ByteDance flywheel

AI PRODUCTSDoubao • Douyin • agentsTOKEN DEMANDpersistent inference loadSYSTEM ENGINEERINGchips • network • schedulerVOLCANO ENGINEexternalises capabilityMORE UTILISATIONcustomers + scale
Scroll horizontally to explore the full diagram →
China AI Infrastructure · chinaaiinfra.netlify.app · See accompanying evidence and review date.

The flywheel is unusually powerful because internal demand can justify infrastructure engineering before external cloud customers arrive; Volcano Engine can then monetise that engineering externally. The key unresolved question is whether this improves total system cost per useful completed task after HBM premiums, additional networking, software complexity and power infrastructure are included.

Evidence ledger

PRIMARY ByteDance Seed infrastructure team: distributed training, high-performance inference, heterogeneous hardware compilation.
PRIMARY ByteDance Seed COMET: production deployment in tens-of-thousands-of-GPU clusters.
PRIMARY Volcano Engine veRoCE: protocol and company-reported performance results.
PRIMARY Volcano Engine EthLink: Ethernet GPU scale-up interconnect design.
PRIMARY Volcano Engine / ByteDance storage: EB-scale HDFS infrastructure.
PRIMARY China Energy Engineering: Wuhu 220kV EPC and Wuwei 220kV EPC.
PRIMARY Inner Mongolia Development and Reform Commission: ByteDance talks on compute monitoring/trading and token economy.
PRIMARY DYXnet: Volcano Engine cooperation in cloud, AI, security and network.
REUTERS Iluvatar / Kunlunxin discussions; HBM shortage and supplier ranking; custom CPUs; Arm data-centre CPU use; reported Qualcomm ASIC deal; $29.6bn financing; Thailand infrastructure expansion.

Research gaps generated by the ecosystem

Physical suppliers

Chinese market commentary names cooling, optics and power suppliers around ByteDance AI Rack projects, but much of it is not primary disclosure. We should not promote those names to confirmed graph edges until customer evidence is stronger.

Third-party IDC exposure

We now have a genuine Volcano Engine–DYXnet/VNET edge, but not evidence that VNET or GDS hosts a material share of ByteDance's AI capacity. This remains a high-value research gap rather than an inference.

Accelerator workload split

Which workloads actually run on Huawei vs Cambricon vs Iluvatar vs Nvidia, and at what utilisation/TCO? Supplier ranking is known; workload allocation is not.

Compute ↔ electricity scheduling

Inner Mongolia discussions are highly relevant, but they do not yet demonstrate hyperscale production workloads dynamically following power availability or price.

Method: primary company/government/infrastructure sources are preferred. Reuters sourcing is treated as corroborated reporting, not company confirmation. Vendor performance claims remain vendor-reported. Architectural inference is labelled separately from established commercial relationships.