Mingxin FX and NVIDIA CMX: The Same Tier, Two Delivery Paths
NVIDIA CMX is the Rubin/BlueField-4 first-party KV context tier; Mingxin FX builds the same tier, already in production, with signed 480B throughput +29–40%.
What CMX solves — stated fairly
NVIDIA CMX (Context Memory eXtension / Context Memory Storage Platform) is the 2026-public AI-native context tier: BlueField-4 manages Ethernet-attached flash, inserting a G3.5 layer between HBM, host RAM, local SSD and traditional network storage, purpose-built for long-context, multi-turn and agentic KV cache (public-source basis: https://www.nvidia.com/en-us/data-center/ai-storage/cmx/).
On the software side, Dynamo does KV-aware routing, NIXL moves the blocks, and DOCA Memos exposes flash as a pod-level KV API; the fabric is Spectrum-X Ethernet. For an all-NVIDIA Rubin cluster this is the first-party co-designed path — something a third-party array cannot structurally match on topology and energy co-design.
Performance: the NVIDIA product page states up to 5× throughput and 5× power efficiency versus traditional storage. The baseline and workload are not disclosed. This site does not cross-compare that claim with Mingxin measurements, and does not repeat other multipliers that are not on the product page.
Availability: the product page does not state a GA date. Partner and industry reports point to 2H 2026 with Rubin / BlueField-4. Until it ships, CMX is not a “joint-test now” option.
Where Mingxin FX sits
Same tier, different binding. The FX all-flash NVMe-oF array plus the vLLM/LMCache stack also sinks KV from HBM into a shareable flash pool. The difference: no BlueField-4 / Rubin lock-in, standard 100 GbE RoCEv2, and source-level adaptation of an open stack (parallel-read patch: R1/R8).
Measured anchors (signed, downloadable): 480B production-form throughput +29–40%, TTFT −26–32%, 8.6–20× vs no-external-storage re-compute (R2/R3); cross-instance fs:// hot sharing verified (R3); platforms AMD MI308X (R1–R4), Ascend 910B (R9), MetaX N260 (R5 methodology).
In one line: on NVIDIA Rubin clusters, CMX is the first-party path and FX can be a standard NVMe-oF sink — interop with the CMX / DOCA Memos stack has not been jointly tested and is not claimed. On non-NVIDIA clusters, no public CMX adaptation exists; FX is the measured, shipping full stack. The control comparison is re-compute, not unshipped CMX.
How to choose
All-NVIDIA Rubin, willing to follow the first-party cadence: evaluate CMX first. Domestic / mixed GPUs (MI308X, Ascend, MetaX), or a gated joint test now: put Mingxin FX on the shortlist. Side-by-side tests are welcome — the G3 gate (TTFT reduction ≥25%, throughput inside the measured +29–40% band) is fair to every vendor.
A one-slide architecture deck is on the Architecture page, for presenting or printing.
FAQ
Is CMX’s 5× faster than Mingxin’s 29–40%?
They are not comparable. 5× is NVIDIA’s vendor figure versus “traditional storage,” with no public baseline; 29–40% is a signed 480B TP8 measurement versus a local NVMe drive (R2/R3). Different denominators.
If I run Dynamo, do I still need CMX or FX?
Yes. Dynamo orchestrates; the KV bytes still land on a medium. CMX is NVIDIA’s first-party G3.5 medium; FX is a standard NVMe-oF medium. Framework and storage tier are complementary.
Can Mingxin speak DOCA Memos / CMX?
Not jointly tested, so not claimed. The published Dynamo/NIXL stance stops at “standard NVMe-oF sink.” First-party-stack interop will be backfilled only with a joint-test report.
Data sources (verifiable)
Related reading
- Mingxin FX and NVIDIA Dynamo: Complementary Roles
- KV Cache Tiering: HBM → RAM → All-Flash Architecture
- KV Cache Offloading: Inference Cache on All-Flash
Joint test first, decisions second: gate-based acceptance with built-in stop-loss
The full costing model is provided as reproducible Python after NDA — customers can rerun it with their own parameters. Every key figure on this site carries a report ID and is open to third-party verification.
This site presents business-cooperation information and constitutes neither an investment offer nor any promise of returns. Measured data come from signed / official test reports (see the Evidence Library); vendor specs, public sources and estimates are labeled as such.