Frequently Asked Questions
Common questions on products, performance data, pricing and how to work with us. Every number cites its report — see theEvidence Library.
QWhat is the Mingxin FX series storage acceleration platform?
The FX series is Mingxin's self-developed all-flash NVMe-oF storage acceleration platform (FX100/FX200/FX300 in production; FX400 GA expected late 2026), paired with a KV cache tiering software stack (vLLM + LMCache, including Mingxin's parallel-read patch) to accelerate LLM inference and training workloads on the storage side. FX100: 100 Gb per port, 16M IOPS; FX300: PCIe 5.0 platform, 400 Gb per port, 60M IOPS.
QHow much inference performance does KV cache tiering add?
In signed benchmarks on a 480B model (Qwen3-Coder-480B-FP8) in production deployment form on 8× AMD MI308X: inference throughput up 29–40% (+29% at concurrency 8, +40% at the optimal operating point), time-to-first-token (TTFT) down 26–32%, and 8.6–20× faster than a no-external-storage re-compute scheme. Sources: test reports R2/R3, downloadable in the Evidence Library for verification.
QHow are these numbers verified — are they reproducible?
Every key figure comes from signed / official test reports (R1–R9; R4 is an official report issued by a third-party testing organization), downloadable directly from the Evidence Library. Test code (the LMCache parallel-read git patch, load clients, orchestration scripts, forensic probes) and raw data are packaged as R8, enabling full independent third-party reproduction. Two independent runs at the same load level deviate by ≤5%.
QHow much does a fully populated FX machine cost?
Fully-populated reference prices (incl. 24 enterprise NVMe SSDs): FX100 ≈ ¥371,200 (≈ ¥2,014/TB); FX200 ≈ ¥331,200 (≈ ¥1,797/TB, the line's lowest per-TB cost); FX300+DPU ≈ ¥924,000 (≈ ¥5,014/TB, the Gen5 performance tier). SSD prices follow the market; final pricing per purchase contract, with a reproducible calculation.
QWe don't want to add GPUs to our existing cluster — does storage acceleration help?
Yes — that is the "efficiency before expansion" route. In measurements against an NFS baseline on Huawei Ascend 910B (R9): DeepSeek-32B serving load fell from 691 s to 112 s (6.2×) and DeepSeek-70B from 1399 s to 150 s (9.3×); training checkpoint saves are 1.9× faster. Effective compute utilization under model switching can rise from 46.7% to 62.8% (R1, at 20 switches/hour).
QWhich domestic / non-NVIDIA accelerators are supported?
Measured platforms include AMD Instinct MI308X (ROCm 7.2, R1–R4), Huawei Ascend 910B (R9) and MetaX N260 (R5 methodology). Mingxin has source-level inference-stack engineering capability (vLLM/LMCache ROCm builds, the parallel-read patch, a GDS adaptation layer); in multimodal, all 7 ComfyUI + LTX-Video 2.3 models ran end-to-end on MI308X (R6/R7).
QHow do we start working together?
Cooperation starts with a gated joint test: about 10 weeks across four gates (G1 delivery acceptance → G2 single-node baseline → G3 main joint-test gate: TTFT reduction ≥25% and throughput inside the measured +29–40% band → G4 72-hour stability acceptance), with stop-loss if any gate is missed — scale discussions follow passing gates. The full costing model is provided as Python after NDA so you can rerun it with your own parameters. Book a slot via the contact page.
QHow do AISSD5000, WS5000 and FX100 relate?
Mingxin FX100 appears in historical test-report filenames as AISSD5000 (also WS5000 / GP5000) — all are names for the same product. This site uses the unified FX naming (FX100/FX200/FX300/FX400); report entries keep original filenames for verification.