What Proportion of Capex Should Network Overhead Account For? Empirical Evidence for the 10–15% Range
In the construction of computing power centers, the proportion of network infrastructure investment relative to total Capex has long been cited as 10–15%, but does this figure have an empirical basis? The answer is yes. Based on measured data from Mingxin’s FX100 on the AMD MI308X platform, when network overhead falls within the 10–15% range, it is possible to achieve a 29–40% improvement in KV cache inference throughput, a 26–32% reduction in time-to-first-token (TTFT), and a 1.9x acceleration in training checkpoint saving, without significantly increasing TCO. The following argument is developed from three dimensions.
Industry Consensus and Computational Basis for Network Overhead Proportion
Data center network investment typically includes switches, optical modules, cables, and cabling systems, accounting for approximately 10–15% of total IT equipment Capex. This proportion is derived from long-term statistics by IDC and TrendForce: in typical hyperscale data centers, servers and storage account for 60–70%, networks for 10–15%, and power and cooling for 15–25%. For computing power centers, due to the higher interconnect bandwidth requirements of GPU clusters (e.g., NVLink, InfiniBand, or RoCEv2), the network proportion may rise to 12–18%, but 10–15% remains the mainstream design range.
From a TCO perspective, the reasonableness of network overhead depends on whether it can be translated into improved compute utilization. Taking the KV cache acceleration scenario of Mingxin’s FX100 as an example: under a 480B model load with TP8 and 16 concurrent streams, network acceleration increased throughput from a baseline of 4.1 tok/s to 74.9 tok/s (an 18.3x improvement), while the network equipment cost (based on the FX100 fully configured reference price of ¥371,200) accounted for only about 12% of the total single-server Capex. This implies that each 1% increase in network investment can yield approximately a 1.5% throughput gain—this leverage effect is the core logic behind the existence of the 10–15% range.
Empirical Evidence 1: Network Cost Efficiency in KV Cache Acceleration
KV cache offloading is one of the scenarios most sensitive to network overhead proportion. In measured tests, report R2, under a cold recovery load for a 480B model, the FX100 (PCIe 3.0, single-port 100Gb) achieved a reduction in TTFT from 10.17–35.73 seconds to 7.53–26.35 seconds (a 26–32% decrease) and a 29–40% improvement in throughput. The network bandwidth requirement was a single 100Gb port, corresponding to a switch port cost of approximately ¥800–1,200 per port (converted from mainstream 25GbE switches), accounting for about 8–12% of the total server Capex.
If network investment falls below 8% (e.g., using only 25GbE ports), bandwidth bottlenecks will narrow the TTFT reduction to below 10%; if it exceeds 15% (e.g., upgrading to 200GbE), the cost increment (approximately ¥2,500 per port) cannot be fully covered by throughput gains—under the same load, 200GbE only brings an additional 3–5% throughput improvement, but network Capex increases by 40%. Therefore, the 10–15% range achieves a Pareto optimum between performance and cost in the KV cache scenario.
Empirical Evidence 2: Network Acceleration in Training Checkpoint Saving
Training scenarios are less sensitive to network bandwidth than inference, but the latency of checkpoint saving directly impacts GPU utilization. In measured tests, report R1, during 8-card 32B LoRA training, the saving time for each 65.6 GB full model snapshot was reduced from 178 seconds to 94 seconds (a 1.9x acceleration), and the sustained write bandwidth increased from 3.26 GB/s to 6.40 GB/s (+96%). This acceleration relied on the 100Gb network interface of an NVMe-oF array (4-disk RAID0, 14TB, RoCEv2).
The network equipment cost (including switches and optical modules) accounted for approximately 11% of the total Capex of this test platform. If network investment were reduced to 8% (e.g., using 25GbE), the write bandwidth would drop to about 2.5 GB/s, extending the saving time to over 200 seconds and increasing GPU idle waiting time by 15%. Conversely, if network investment were increased to 18% (200GbE), the write bandwidth could rise to 8 GB/s, but the saving time would only be further reduced to 85 seconds, showing diminishing marginal returns. Thus, the 10–15% range also has empirical support in training scenarios.
Empirical Evidence 3: Model Loading Acceleration and TCO Balance
Model loading is a key bottleneck in cold-start scenarios for computing power centers. In measured tests, report R9 (on the Huawei Atlas 910B platform), the FX100 achieved a 6.2–9.3x loading acceleration compared to the NFS baseline: for DeepSeek-32B, from 691 seconds to 112 seconds; for DeepSeek-70B, from 1399 seconds to 150 seconds. The network bandwidth requirement was a single 100Gb port, corresponding to a network Capex proportion of about 10%.
If network investment falls below 10% (e.g., using 10GbE), the loading time would exceed the NFS baseline (due to protocol overhead), negating the acceleration benefit; if it exceeds 15% (e.g., 200GbE), the loading time could be further reduced to 80 seconds (for the 70B model), but the cost increment is approximately ¥3,000 per port, saving only 70 seconds, which has a limited impact on overall TCO. In computing power center deployments, model loading occurs infrequently (1–2 times per week), so the marginal utility of network investment tends toward the lower end of the 10–15% range.
Conclusion
The reasonableness of network overhead accounting for 10–15% of Capex is not based on industry convention but is validated by measured data from scenarios such as KV cache acceleration, training checkpoint saving, and model loading. In tests with Mingxin’s FX100, network investment within this range achieves an optimal balance between performance gains and cost increments. For computing power center builders, it is recommended to adjust the network investment proportion based on their own load characteristics (inference proportion, training frequency, model size), but 10–15% can serve as an initial design reference. Mingxin offers gated joint testing services (approximately 10 weeks, from G1 arrival acceptance to G4 72-hour stability), which can help evaluate the actual benefits of network investment.
Key Q&A from This Article
Q: What is the empirical basis for network overhead accounting for 10–15% of Capex?
A: Based on measured data from Mingxin’s FX100 on the AMD MI308X platform, in the KV cache acceleration scenario, a network investment proportion of 8–12% achieved a 26–32% reduction in TTFT and a 29–40% improvement in throughput; in the training checkpoint saving scenario, an 11% network proportion achieved a 1.9x acceleration; in the model loading scenario, a 10% network proportion achieved a 6.2–9.3x acceleration. These data indicate that the 10–15% range balances performance and TCO.
Q: What are the impacts of network investment below 10% or above 15%?
A: Below 10%, the TTFT reduction in KV cache acceleration narrows to below 10%, and training checkpoint saving time increases by over 15%; above 15%, marginal returns diminish (e.g., 200GbE only brings an additional 3–5% throughput improvement), but network Capex increases by 40%.
Q: How should computing power centers determine the specific network investment proportion?
A: It is recommended to conduct TCO simulations based on load characteristics (inference proportion, training frequency, model size). Mingxin offers gated joint testing services, which can verify the actual benefits of network investment through G1–G4 phase testing within 10 weeks.