Mingxin

Spare Parts Rate, Maintenance Fees, Installation Costs: Hidden Costs Easily Underestimated in Compute Centers

算力中心TCO数据中心算力建设

In the TCO (Total Cost of Ownership) of a compute center, hardware procurement is only part of the visible costs. The three hidden expenditures—spare parts rate, maintenance fees, and installation costs—can account for 18%–30% of hardware procurement over a three-year lifecycle, with volatility far exceeding that of the equipment itself. Based on measured data from the Mingxin FX100 on the AMD MI308X platform, this article provides a quantitative breakdown and decision-making recommendations for these three cost items.

Spare Parts Rate: Cash Flow Pressure Behind Redundant Design

The spare parts rate is typically provisioned at 5%–15% of hardware procurement, but most compute centers plan only at the minimum standard. Taking the Mingxin FX100 all-flash NVMe-oF array as an example, the fully configured reference price is ¥371,200 (approximately ¥2,014/TB). At a 10% spare parts rate, an additional ¥37,120 per unit must be reserved. For a thousand-GPU cluster, this figure scales to millions.

The key issue is whether the provisioning standard for the spare parts rate matches the actual failure rate. In R2/R3 testing, the Mingxin FX100 ran continuously for 72 hours (G4 gate) under a 480B model long-context cold restart workload without storage-side failures. However, the difference between test and production environments lies in the frequent model hot-switching, multi-tenant concurrency, and uncontrollable operational actions in production. It is recommended to tier the spare parts rate by equipment category—storage arrays at 8%–10%, network equipment at 5%–8%, and compute nodes at 3%–5%. For NVMe-oF-based storage, SSD wear leveling and lifespan management should be included in spare parts assessment, not just capacity.

Maintenance Fees: Annual Service Premium

Maintenance fees are typically 8%–15% of hardware procurement per year, but vary significantly by service level agreement (SLA). In R1 testing, the Mingxin FX100 delivered 6.2–9.3× acceleration in model inference loading on an 8-GPU MI308X platform (vs NFS baseline, measured, report R9), which relies on the continuous healthy operation of the storage system. If maintenance services cover only hardware replacement without firmware upgrades and performance tuning, the actual value is greatly diminished.

A frequently overlooked dimension is whether maintenance fees include performance regression testing. In R2/R3 testing, the Mingxin FX100's KV Cache acceleration reduced TTFT by 26%–32% (480B·TP8 three-tier concurrency, measured, report R2), and such performance metrics fluctuate with firmware versions and driver upgrades. It is recommended to explicitly stipulate in the maintenance contract: a quarterly performance baseline retest to ensure acceleration effects do not degrade. Otherwise, maintenance fees are merely "buying insurance," not "buying performance."

Installation Costs: One-Time Investment, Long-Term Impact

Installation costs may appear to be a one-time expense, but they affect all subsequent operational costs. The Mingxin FX100 supports PCIe 3.0/4.0/5.0/6.0 interfaces (FX100/FX200/FX300/FX400), and installation complexity varies significantly across generations—PCIe 5.0/6.0 demands higher signal integrity, with labor costs for cabling, cooling, and firmware configuration approximately 1.5–2 times that of PCIe 3.0.

More critically, installation quality impacts performance. In R2 testing, the FX100 achieved a throughput increase of 35–36% at the full-machine TP4×2 level (measured, report R3), contingent on correct topology configuration of the storage network (RoCEv2) and compute nodes. If RDMA latency tuning for NVMe-oF is not performed during installation, actual throughput may be discounted by more than 20%. It is recommended to bind installation acceptance to performance gates—in Mingxin's collaboration model, the G3 primary gate requires a TTFT reduction of ≥25% and throughput +29–40% measured in-band, which is essentially a form of "installation quality insurance."

Measurement Framework: Hidden Cost Model for Three-Year TCO

Based on the above analysis, a reproducible measurement framework is provided:

Three-Year TCO = Hardware Procurement × (1 + Spare Parts Rate + Maintenance Fee × 3 + Installation Cost Ratio)

Using the FX100 fully configured unit at ¥371,200 as an example:

  • Spare parts rate at 10%: ¥37,120
  • Maintenance fee at 12%/year × 3 years: ¥133,632
  • Installation cost (including tuning) at 5%: ¥18,560
  • Total hidden costs: ¥189,312, accounting for 51% of hardware procurement

This means that if budgeting is based solely on hardware quotes, actual expenditure will exceed the budget by half. For large-scale compute centers, this difference is sufficient to impact the return on investment model.

Key Q&A

Q: In compute center TCO, what proportion of hardware procurement do spare parts rate, maintenance fees, and installation costs typically account for? A: The spare parts rate is typically provisioned at 5%–15%, maintenance fees at 8%–15%/year, and installation costs at approximately 3%–5%. Over a three-year cycle, the combined total may reach 30%–50% of hardware procurement, representing hidden costs that are often underestimated.

Q: How can one verify that maintenance fees are worth the value? A: It is recommended to stipulate quarterly performance baseline retests in the maintenance contract. Taking the Mingxin FX100 as an example, its KV Cache acceleration achieved a TTFT reduction of 26%–32% and throughput improvement of 29%–40% in R2/R3 testing. If these metrics degrade after maintenance, it indicates that the service does not cover performance upkeep.

Q: How does installation quality affect long-term performance? A: Network topology and RDMA tuning during installation directly impact storage acceleration effectiveness. The Mingxin FX100 achieved a throughput increase of 35–36% at the full-machine TP4×2 level in R2 testing (measured, report R3), contingent on correct RoCEv2 configuration. It is recommended to bind installation acceptance to performance gates, such as Mingxin's G3 gate requiring a TTFT reduction of ≥25%.


Mingxin (Tianjin) Semiconductor Equipment Co., Ltd. provides storage acceleration and full-chain services for compute centers, supporting approximately 10 weeks of gated joint testing (from G1 arrival acceptance to G4 72-hour stability), with stop-loss measures if targets are not met. For TCO calculations based on actual workloads, a Python-reproducible model can be provided under NDA.

Generated by Mingxin's content engine with automated QC; headline numbers cite signed test reports (see the evidence library). Translated from the Chinese original. Questions or corrections: contact us.