Mingxin
MingxinArchitecture
中文

TECHNICAL BRIEF · VS NVIDIA CMX

Same tier.Two paths.

KV cache that leaves HBM should not be recomputed. NVIDIA CMX and Mingxin FX fill the same G3.5 context-storage tier.

480B long-context cold recovery · TTFT p50 · concurrency 16

149.5sRecompute, no external store
11.85sFX100 tiering

8.6–20× vs recompute · R2 signed

G1
HBMHot sessions192 GB / GPU
G2
Host RAMRecent≈1.5 TB / node
G3
Local NVMePer node2 TB single drive
CMX or FX lives here
G3.5
Context storageShared KV across nodes184 TB / unit (FX100 full)
G4
Network storageModels & archiveNFS / object
Faster · smaller · costlier

Both fill G3.5. Neither replaces HBM. · Capacities are the R2 test-platform configuration

Larger · slower · shareable

Same tier. Different binding.

5× and +29–40% use different denominators. Do not cross-compare.

NVIDIA CMX

Public / vendor
Up to 5×vs “traditional storage” · baseline undisclosed
Binding
Rubin · BlueField-4
GPUs
NVIDIA accelerators
Stack
Dynamo · NIXL · DOCA Memos
Ship
Reported 2H 2026

Mingxin FX

Signed
+29–40%480B TP8 throughput · TTFT −26–32% · R2/R3
Binding
Standard 100 GbE RoCEv2
GPUs
MI308X · Ascend · MetaX
Stack
vLLM · LMCache, source-level
Ship
FX100 / 200 / 300 shipping
01
All-NVIDIA RubinEvaluate first-party CMX first
02
Mixed / domestic GPUsFX ships today, measured
03
Need a test nowTen-week gates, stop-loss if missed

NVIDIA figures are vendor-public with no disclosed baseline. Every Mingxin figure carries a report ID (R1–R9) and is open to re-verification. Not a disparagement or substitution claim. NVIDIA CMX product page · NVIDIA Technical Blog