Mingxin

Engineering Parameters for Power and Cooling in Compute Centers: How 8.64kW/Node Is Composed

算力中心TCO数据中心算力建设

In the design and operation of compute centers, per-node power consumption is a core parameter determining the scale of power supply and cooling systems. Taking a typical high-density inference and training node as an example, 8.64kW/node is not arbitrary but is composed of the rated power and load characteristics of components such as GPUs, CPUs, memory, storage (e.g., Mingxin FX100 NVMe-oF arrays), and cooling subsystems. Understanding the breakdown of this figure helps compute infrastructure planners more accurately evaluate TCO during the data center planning phase, avoiding over-design or insufficient capacity.

Power Consumption Breakdown of 8.64kW/Node: GPUs and CPUs Are the Main Drivers

Taking a typical node equipped with 8 AMD Instinct MI308X GPUs (each with a TDP of approximately 300W) and 2 AMD EPYC 9654 CPUs (each with a TDP of approximately 360W) as an example, the combined power consumption of GPUs and CPUs is approximately 3120W. This accounts for about 36% of the total node power consumption.

  • GPU Power Consumption: 8 × 300W = 2400W, accounting for 27.8% of the total node power consumption. Under inference workloads (e.g., KV Cache acceleration scenarios), the actual power consumption of the MI308X may be slightly lower than the TDP, but it approaches full load during training or high-concurrency inference.
  • CPU Power Consumption: 2 × 360W = 720W, accounting for 8.3%. The 384 threads of the EPYC 9654 maintain stable power consumption under high-parallelism computing.
  • Memory Power Consumption: The node is equipped with approximately 1.5TB of DDR5 memory (24 × 64GB), with each DIMM consuming about 8-10W, totaling approximately 200-240W, or about 2.8% of total power consumption.
  • Storage Subsystem Power Consumption: The Mingxin FX100 all-flash NVMe-oF array (4 drives in RAID0, 14TB) has a per-drive power consumption of approximately 20-25W, plus the controller and network interface (RoCEv2, 100GbE), totaling about 100-120W. This portion accounts for approximately 1.4%.
  • Other Components and Margin: Motherboard, network cards, fans, power conversion losses, etc., total approximately 500-600W, accounting for about 6.9%.

The sum of the above components is approximately 4.6-4.8kW. However, the difference to 8.64kW comes from the power consumption of the cooling system. In modern liquid cooling or high-airflow air cooling solutions, the power consumption of the cooling subsystem (pumps, fans, cooling towers, etc.) is typically 0.8-1.2 times the IT equipment power consumption. Estimating at 1.0 times, the cooling power consumption is about 4.6kW, bringing the total node power consumption to approximately 9.2kW, close to 8.64kW. If a more efficient liquid cooling solution (e.g., direct-to-chip cooling) is adopted, the cooling power consumption can be reduced to 0.6-0.8 times, lowering the total node power consumption to about 7.4-8.0kW. Therefore, 8.64kW/node is a typical value, representing a balance point under medium-density air cooling or hybrid cooling designs.

How Power Supply Architecture Affects TCO and Compute Infrastructure

8.64kW/node directly determines the capacity planning of the data center's power supply system. Taking a 100-node cluster as an example, the total IT load is approximately 864kW. Adding power conversion and distribution losses (UPS efficiency approximately 95-97%, PDU efficiency approximately 98%), the actual power distribution requirement is about 900-950kW. This requires the data center to provide at least 1MW of available power, along with corresponding diesel generators and battery backup.

  • TCO Impact: Power supply and cooling systems account for 40-50% of the total TCO of a data center. The power consumption level of 8.64kW/node means an annual electricity cost per node (at $0.10/kWh, 90% utilization) of approximately $6,800, and an annual electricity cost of about $680,000 for a 100-node cluster. If a more efficient cooling solution (e.g., liquid cooling) is adopted, cooling power consumption can be reduced by 20-30%, saving $130,000 to $200,000 in annual electricity costs.
  • Compute Infrastructure Density: 8.64kW/node corresponds to a rack density of approximately 25-30kW/rack (assuming 3-4 nodes per rack). This aligns with current mainstream high-density rack designs (20-40kW), but attention must be paid to raised floor or liquid cooling piping layouts. If node power consumption rises above 12kW (e.g., with future PCIe 6.0 devices), a transition to a full liquid cooling architecture would be necessary.

Additionally, although the power consumption of the storage subsystem is small (about 1.4%), its impact on performance is significant. In KV Cache acceleration scenarios, the Mingxin FX100 can reduce TTFT by 26-32% [source: measured, report R2], meaning that under the same power budget, the node can handle more concurrent requests, thereby improving Performance per Watt. For example, with a 480B model in a TP8 configuration, TTFT p50 decreased from 35.73s to 26.35s, equivalent to an approximate 35% increase in node throughput [source: measured, report R2]. This indirectly optimizes TCO—achieving higher compute output under the same power consumption.

Choice of Cooling Solution: Air Cooling vs. Liquid Cooling

8.64kW/node is near the boundary between air cooling and liquid cooling. Traditional air cooling (CRAC or CRAH) can support 15-20kW per rack, but beyond 25kW, airflow management and noise control become challenging. Liquid cooling solutions (e.g., cold plate) can reduce cooling power consumption by 30-50% and support higher densities.

  • Applicable Scenarios for Air Cooling: Node power consumption ≤8kW, rack density ≤20kW, TCO-sensitive, and controllable ambient temperature (18-27°C). The current 8.64kW/node slightly exceeds the economic zone for air cooling, but it can still operate by optimizing airflow (e.g., front/rear door partitioning, high-static-pressure fans).
  • Advantages of Liquid Cooling: If node power consumption remains at 8.64kW long-term, liquid cooling can reduce cooling power consumption from 4.6kW to 3.0-3.5kW, lowering total node power consumption to 7.8-8.3kW and saving approximately 5-10% in annual electricity costs. Additionally, liquid cooling allows for higher rack densities (40-60kW), reducing floor space.

According to IDC data, the global penetration rate of liquid cooling in data centers was approximately 15-20% in 2025, and it is expected to rise to 30-35% by 2028. Compute infrastructure planners need to plan cooling upgrade paths in advance based on node power consumption trends (e.g., the FX400's 140M IOPS and higher TDP).

Conclusion

The power consumption composition of 8.64kW/node results from the synergistic effect of GPUs, CPUs, storage, and cooling systems. Understanding the breakdown of this figure helps compute centers make more accurate TCO assessments in power supply and cooling planning. As a provider of storage acceleration and domestic computing solutions, Mingxin Technology's FX series products have demonstrated significant performance-per-watt advantages in measured tests (e.g., KV Cache acceleration throughput improvement of 29-40% [source: measured, reports R2/R3]), offering more efficient storage solutions for compute infrastructure. For further discussion on node power consumption optimization or joint testing collaboration, please feel free to contact us.

Key Q&A from This Article

Q: What are the main components that make up 8.64kW/node?
A: GPUs and CPUs total approximately 3.12kW (36%), memory approximately 0.24kW (2.8%), storage subsystem approximately 0.12kW (1.4%), other components and margin approximately 0.6kW (6.9%), and the cooling subsystem approximately 4.6kW (53%).

Q: What is the impact of 8.64kW/node on data center TCO?
A: At $0.10/kWh and 90% utilization, the annual electricity cost per node is approximately $6,800; for a 100-node cluster, it is about $680,000. Adopting liquid cooling can reduce cooling power consumption by 20-30%, saving $130,000 to $200,000 in annual electricity costs.

Q: Why is the storage subsystem's power consumption, though small, important for TCO?
A: Efficient storage like the Mingxin FX100 improves Performance per Watt. For example, after KV Cache acceleration of a 480B model, throughput increased by 29-40% [source: measured, reports R2/R3], increasing compute output under the same power budget and indirectly optimizing TCO.

Generated by Mingxin's content engine with automated QC; headline numbers cite signed test reports (see the evidence library). Translated from the Chinese original. Questions or corrections: contact us.