Mingxin

Mingxin FX and NVIDIA CMX: The Same Tier, Two Delivery Paths

PublishedUpdated
Direct answer

NVIDIA CMX is the Rubin/BlueField-4 first-party KV context tier; Mingxin FX builds the same tier, already in production, with signed 480B throughput +29–40%.

What CMX solves — stated fairly

Key point

NVIDIA CMX (Context Memory eXtension / Context Memory Storage Platform) is the 2026-public AI-native context tier: BlueField-4 manages Ethernet-attached flash, inserting a G3.5 layer between HBM, host RAM, local SSD and traditional network storage, purpose-built for long-context, multi-turn and agentic KV cache (public-source basis: https://www.nvidia.com/en-us/data-center/ai-storage/cmx/).

On the software side, Dynamo does KV-aware routing, NIXL moves the blocks, and DOCA Memos exposes flash as a pod-level KV API; the fabric is Spectrum-X Ethernet. For an all-NVIDIA Rubin cluster this is the first-party co-designed path — something a third-party array cannot structurally match on topology and energy co-design.

Performance: the NVIDIA product page states up to 5× throughput and 5× power efficiency versus traditional storage. The baseline and workload are not disclosed. This site does not cross-compare that claim with Mingxin measurements, and does not repeat other multipliers that are not on the product page.

Availability: the product page does not state a GA date. Partner and industry reports point to 2H 2026 with Rubin / BlueField-4. Until it ships, CMX is not a “joint-test now” option.

Where Mingxin FX sits

Key point

Same tier, different binding. The FX all-flash NVMe-oF array plus the vLLM/LMCache stack also sinks KV from HBM into a shareable flash pool. The difference: no BlueField-4 / Rubin lock-in, standard 100 GbE RoCEv2, and source-level adaptation of an open stack (parallel-read patch: R1/R8).

Measured anchors (signed, downloadable): 480B production-form throughput +29–40%, TTFT −26–32%, 8.6–20× vs no-external-storage re-compute (R2/R3); cross-instance fs:// hot sharing verified (R3); platforms AMD MI308X (R1–R4), Ascend 910B (R9), MetaX N260 (R5 methodology).

In one line: on NVIDIA Rubin clusters, CMX is the first-party path and FX can be a standard NVMe-oF sink — interop with the CMX / DOCA Memos stack has not been jointly tested and is not claimed. On non-NVIDIA clusters, no public CMX adaptation exists; FX is the measured, shipping full stack. The control comparison is re-compute, not unshipped CMX.

How to choose

Key point

All-NVIDIA Rubin, willing to follow the first-party cadence: evaluate CMX first. Domestic / mixed GPUs (MI308X, Ascend, MetaX), or a gated joint test now: put Mingxin FX on the shortlist. Side-by-side tests are welcome — the G3 gate (TTFT reduction ≥25%, throughput inside the measured +29–40% band) is fair to every vendor.

A one-slide architecture deck is on the Architecture page, for presenting or printing.

FAQ

Is CMX’s 5× faster than Mingxin’s 29–40%?

They are not comparable. 5× is NVIDIA’s vendor figure versus “traditional storage,” with no public baseline; 29–40% is a signed 480B TP8 measurement versus a local NVMe drive (R2/R3). Different denominators.

If I run Dynamo, do I still need CMX or FX?

Yes. Dynamo orchestrates; the KV bytes still land on a medium. CMX is NVIDIA’s first-party G3.5 medium; FX is a standard NVMe-oF medium. Framework and storage tier are complementary.

Can Mingxin speak DOCA Memos / CMX?

Not jointly tested, so not claimed. The published Dynamo/NIXL stance stops at “standard NVMe-oF sink.” First-party-stack interop will be backfilled only with a joint-test report.

Data sources (verifiable)

R1FX100 Comprehensive LLM Inference & Training Benchmark (8× AMD MI308X)2026-07-03
Download report PDF ↓
R2FX100 KV-Cache Benchmark (480B, TP8 long-context, signed)2026-07-05
Download report PDF ↓
R3FX100 KV-Cache Benchmark Summary (480B, TP4×2, all metrics, brand-unified edition)2026-07-06
Download report PDF ↓
R8FX100 KV-Cache AMD Code Export + Raw Benchmark Data2026-07
Contact us for access →

Related reading

Joint test first, decisions second: gate-based acceptance with built-in stop-loss

The full costing model is provided as reproducible Python after NDA — customers can rerun it with their own parameters. Every key figure on this site carries a report ID and is open to third-party verification.

This site presents business-cooperation information and constitutes neither an investment offer nor any promise of returns. Measured data come from signed / official test reports (see the Evidence Library); vendor specs, public sources and estimates are labeled as such.