CHINA AI
INFRASTRUCTURE
Home / Deep dives / National orchestration & control plane
← BACK TO SYSTEM
VALUE-CHAIN DEEP DIVE • SOFTWARE & MANAGEMENT

CONTROL PLANE

The software and management layer that converts fragmented physical infrastructure into useful AI output. The research focus is not who has a dashboard; it is who can observe resources, make allocation decisions, recover failures, route workloads and ultimately monetise orchestration.

The five control layers

AI / MODEL CONTROL
WORKLOAD CONTROL
COMPUTE CONTROL
RESOURCE / NETWORK CONTROL
AIDC + ENERGY CONTROL
Core distinction: internal orchestration is already widespread. Cross-provider, cross-architecture and cross-region orchestration is the more consequential opportunity — and is much less proven.

Current actors

Kingsoft Cloud

StarFlow provides full-process training/inference management, heterogeneous-resource scheduling, topology-aware RDMA placement, observability and GPU fault self-healing. KC has therefore moved beyond raw compute rental into workload and compute management.

VNET

Smart Navigation is a facility-level “O&M brain” deployed across >90% of self-built sites, combining intelligent scheduling, cloud collaboration and autonomous operations. This creates a lower control-plane asset spanning physical capacity, cooling and energy.

Huawei

Agentic Infra combines token production, context memory, runtime and unified general/AI scheduling. Its architecture increasingly joins physical and software planes.

ByteDance / Volcano Engine

Seed develops distributed training, high-performance inference and heterogeneous-hardware compilation. Volcano Engine exposes GPU and mGPU scheduling, topology and sharing controls.

Alibaba Cloud

A vertically integrated benchmark across models, cloud, workload scheduling, compute, storage and networking. It demonstrates what a full-stack control plane can look like when demand and infrastructure sit inside one platform.

National / telecom layer

MIIT's 1+M+N architecture explicitly calls for resource aggregation, selection, monitoring and interconnection scheduling across different subjects, architectures and regions. This is policy-backed market infrastructure, not yet proof of frictionless compute fungibility.

Capability matrix

ActorWorkloadComputeNetwork / storageAIDCEnergyCross-provider?
KCStrongStrongRDMA/storage integrationLimitedLimitedNot established
VNETLimitedEmerging / investigateNetwork + facility visibilityStrongStrong facility-levelNot established
HuaweiStrongStrongStrongArchitecture layerDigital Power / emerging integrationWithin Huawei ecosystem strong; broader portability unresolved
ByteDanceStrongStrongStrongOwn campusesPhysical power footprintNot established
AlibabaStrongStrongStrongStrongInvestigateNot established beyond its cloud ecosystem
1+M+N fabricMarket matchingResource selectionPaths + monitoringIndirectIndirectExplicit policy objective

Where the value could migrate

Utilisation economics

Higher utilisation turns sunk accelerator capex into more billable output. Scheduling, sharing and failure recovery can therefore create value without adding silicon.

Abstraction economics

If software hides enough hardware differences, buyers can choose among more compute pools. That can reduce dependence on one accelerator source while increasing the value of the abstraction layer.

Neutral orchestration

The largest white space is a trusted layer above individual providers that can discover, compare and route work across operators, regions and architectures. Policy is building some prerequisites; commercial ownership is unresolved.

Physical control

VNET shows why the control plane extends below Kubernetes. AI-ready capacity depends on cooling, power, maintenance, reliability and site-level decisions.

Energy control

CATL/DeepCtrls and VNET's energy-management work suggest a future boundary where compute scheduling and physical energy optimisation interact. Hyperscale dynamic coordination remains unproven.

Failure mode

If every provider remains a closed island, accelerator portability stays poor and national scheduling is mostly directory/trading infrastructure, the control-plane value pool will remain fragmented rather than becoming a dominant neutral layer.

Update · national coordination layer moves toward technical specification

21 September 2026: The National Data Administration standards programme now includes compute-grid connection, monitoring interfaces, resource management/scheduling, multidimensional billing, operations and matching transactions, plus compute–electricity coordination and data-centre adjustable-load assessment. This is unusually close to the neutral orchestration layer this page has been testing. It does not establish that cross-cloud workloads are already fungible or that a single national operator controls scheduling; it does establish that common technical and commercial interfaces are being specified.

The control-plane thesis should therefore be split into two interacting layers: provider control planes (Huawei, Volcano Engine, Alibaba, KC, telecoms and others) and an emerging national coordination layer concerned with discovery, identification, monitoring, scheduling, billing/trading and cross-region resource management. The unresolved value question is which functions remain public/common infrastructure and which become monetisable services for cloud, telecom and neutral infrastructure operators.

National standards register — adjustable-load project and related standards

Research priorities

1. Map exact accelerator support under KC StarFlow and Galaxy Stack. 2. Determine whether VNET exposes Smart Navigation capabilities to customers or only operates its own estate. 3. Identify regional 1+M+N operators and commercial platforms. 4. Trace telecom compute-network scheduling products. 5. Compare Huawei/Alibaba/Volcano/KC APIs and resource models. 6. Find evidence of real cross-cloud workload movement rather than resource discovery alone. 7. Test whether energy signals ever enter workload scheduling decisions.

Primary evidence

Kingsoft Cloud — StarFlow
Kingsoft Cloud — cloud-native AI suite
VNET — Innovation / Smart Navigation
Huawei — Agentic Infra
ByteDance Seed — Infrastructures
Volcano Engine — GPU scheduling
MIIT — national compute-interconnection nodes
MIIT — Compute Interconnection Action Plan