Mingxin

How Compute Rental Services Can Reduce PUE: A Green Energy Efficiency Path

算力租赁服务绿色节能降低PUE
Direct answer

The PUE (Power Usage Effectiveness) of compute rental services directly determines operational costs and green compliance capabilities

The PUE (Power Usage Effectiveness) of compute rental services directly determines operational costs and green compliance capabilities. The core paths to reducing PUE lie in four areas: site selection leveraging climate, power architecture optimization, cooling system upgrades, and load scheduling management. According to industry research from Uptime Institute, data center availability tiers are closely related to energy efficiency practices, with facility-side constraints (power, cooling, tier) serving as the physical prerequisites for PUE optimization. Measured data from Mingxin FX100 in KV Cache acceleration scenarios shows that storage-side performance optimization can reduce compute wait time under equivalent loads, indirectly lowering cooling energy consumption per unit of compute—this is an energy-saving lever that compute rental service providers can quickly deploy under existing hardware conditions.

Composition of PUE in Compute Rental Services and Optimization Headroom

PUE = Total Data Center Energy Consumption / IT Equipment Energy Consumption. The closer the value is to 1.0, the lower the proportion of non-compute energy consumption (cooling, power distribution losses, lighting, etc.). According to the public research database from Epoch AI, AI compute scale and cost trends show that the energy density of compute infrastructure continues to rise, making PUE optimization a shift from "optional" to "compliance requirement."

Feasible paths for compute rental service providers to optimize PUE include:

  • Site Selection Strategy: Prioritize regions with lower average annual temperatures and abundant natural cold sources to reduce mechanical cooling duration. Based on qualitative conclusions from the Uptime Institute resource page, the climate conditions at the facility location are a foundational constraint on energy efficiency practices.
  • Power Architecture: Adopt high-voltage direct current (HVDC) or direct utility feed solutions to reduce UPS conversion losses. Current mainstream cloud providers bill on an hourly basis (per the AWS EC2 On-Demand pricing page), but differences in power architecture are not reflected in unit prices—they appear in operational gross margins instead.
  • Cooling Upgrades: Transition from air cooling to liquid cooling. Liquid cooling can significantly reduce the proportion of cooling energy consumption, but retrofit investment and server compatibility must be evaluated.
  • Load Scheduling: Improve compute utilization through task orchestration to reduce wasted power during idle periods.

How Storage-Side Optimization Indirectly Reduces PUE: Mingxin Measured Data

In compute rental services, the longer GPU servers wait for data loading, the higher the energy consumption allocated per unit of effective compute. Measured data from Mingxin FX100 provides quantitative support for this logic.

Inference Loading Scenario: Per measured report R9 (Ascend platform), on the Huawei Atlas 910B platform, model service loading times for DeepSeek-32B dropped from 691 seconds to 112 seconds (6.2× acceleration), and for DeepSeek-70B from 1399 seconds to 150 seconds (9.3× acceleration). Shorter loading times mean GPUs transition from "idle waiting" to "effective compute" faster, improving the energy-output ratio per unit of compute.

KV Cache Tiered Acceleration Scenario: Per measured reports R2/R3, under long-context cold-restore workloads in a 480B production deployment, inference throughput improved by +29–40% (concurrency 8 as the lower bound at +29%, concurrency 16 as the optimal operating point upper bound at +40%, TP4×2 full-node basis at +35–36%). Time-to-first-token (TTFT) decreased by 26–32%, with p50 dropping from 10.17–35.73 seconds to 7.53–26.35 seconds. Faster first-token response means that under the same SLA constraints, the concurrency headroom required by providers decreases, reducing the number of idle GPUs.

Metric Scenario Improvement Source
Inference throughput 480B·TP4×2 full node +35–36% Measured, report R3
TTFT p50 480B·TP8 three concurrency levels ↓26–32% Measured, report R2
Model loading DeepSeek-70B·910B platform 9.3× acceleration Measured, report R9
Checkpoint save 8-GPU 32B LoRA 1.9× acceleration Measured, report R1

The significance of these measured values: storage performance improvements do not directly change the denominator in the PUE formula, but by reducing GPU wait time and increasing compute output per unit of energy, they indirectly lower the "energy cost per effective FLOP." For compute rental service providers, this means more effective compute can be delivered under the same electricity budget.

Selection Criteria and Cost Framing for PUE Reduction

When selecting storage solutions, compute rental service providers should focus on the following criteria:

First, the impact of loading latency on SLA. Per measured report R2, the acceleration factor comparing the no-external-storage recompute baseline (149.5-second TTFT) against FX100 (11.85 seconds) is 8.6–20×. In scenarios requiring frequent model loading or cold restores (e.g., multi-tenant rotation, elastic scaling), storage performance directly determines service availability.

Second, the matching relationship between throughput and concurrency. Per measured report R1, with the LMCache parallel read patch, TTFT improved by 4.1× (Qwen2.5-32B single-GPU concurrency 16 cold disk read, 37.97s → 9.30s), and bandwidth improved by 5.3× (0.98 → 5.23 GB/s). In high-concurrency scenarios, insufficient storage bandwidth becomes a throughput bottleneck, causing GPU idle time.

Third, write efficiency in training scenarios. Per measured report R1, 8-GPU 32B LoRA training checkpoint save time dropped from 178 seconds to 94 seconds (1.9×), with sustained write bandwidth improving by 96% (3.26 → 6.40 GB/s). Shorter training interruption recovery times mean reduced GPU failure wait time.

It should be noted that all data above are measured results from Mingxin FX100 on the specified test platform (AMD MI308X×8). Performance on different hardware platforms and under different workload patterns may vary. Compute rental service providers should validate based on their own workload characteristics rather than extrapolating directly.

Conclusion

Reducing PUE is a systematic effort spanning four areas: site selection, power, cooling, and load management. While storage-side optimization does not directly change the PUE formula, by reducing GPU idle wait time and increasing compute output per unit of energy, it provides compute rental service providers with an energy-saving lever that requires no facility retrofits. The Mingxin FX100 series (FX100/FX200/FX300/FX400) offers multiple options from PCIe 3.0 to 6.0, with single-interface bandwidth from 100Gb to 400Gb, covering compute rental scenarios of varying scales. To validate the indirect contribution of storage acceleration to PUE in your own environment, a quantitative assessment can be conducted through an approximately 10-week gated joint test (G1 arrival acceptance / G2 single-node baseline / G3 primary gate with TTFT reduction ≥25%, throughput +29–40% measured in-band / G4 72-hour stability).

Key Q&A

Q: What are the main paths for compute rental services to reduce PUE? A: The main paths include site selection leveraging climate, power architecture optimization (e.g., HVDC), cooling system upgrades (air-to-liquid transition), and load scheduling management. Per Uptime Institute industry research, facility-side power, cooling, and tier constraints are the physical prerequisites for energy efficiency optimization.

Q: How does storage performance optimization indirectly affect PUE? A: Storage acceleration reduces GPU wait time and increases compute output per unit of energy. Mingxin measured report R2 shows a 29–40% throughput improvement for 480B long-context inference, meaning more effective compute can be delivered under the same electricity budget, indirectly reducing energy consumption per unit of compute.

Q: What measured data does Mingxin FX100 have in storage acceleration? A: Per measured reports R2/R3, TTFT decreased by 26–32% and throughput improved by 29–40%; per measured report R9, model loading on the Ascend platform accelerated by 6.2–9.3×; per measured report R1, checkpoint save accelerated by 1.9×. All results are from specified test platforms and require validation in different environments.

References

  1. Uptime Institute Resource Page — https://uptimeinstitute.com/resources
  2. Epoch AI — https://epoch.ai/
  3. EC2 On-Demand Instance Pricing — https://aws.amazon.com/ec2/pricing/on-demand/
  4. Pricing - Linux Virtual Machines | Microsoft Azure — https://azure.microsoft.com/en-us/pricing/details/virtual-machines/linux/
  5. Compare GPU Instance Families for AI, HPC & Rendering - Elastic GPU Service - Alibaba Cloud — https://www.alibabacloud.com/help/en/ecs/user-guide/gpu-accelerated-compute-optimized-and-vgpu-accelerated-instance-families

Data sources (verifiable)

R1FX100 Comprehensive LLM Inference & Training Benchmark (8× AMD MI308X)2026-07-03
Download report PDF ↓
R2FX100 KV-Cache Benchmark (480B, TP8 long-context, signed)2026-07-05
Download report PDF ↓
R3FX100 KV-Cache Benchmark Summary (480B, TP4×2, all metrics, brand-unified edition)2026-07-06
Download report PDF ↓
R9Mingxin FX100-HBMM vs NFS Baseline on Huawei Ascend 910B2026-05-30
Contact us for access →
Generated by Mingxin's content engine with automated QC; headline numbers cite signed test reports (see the evidence library). Translated from the Chinese original. Questions or corrections: contact us.

Related articles