Mingxin

Technical Videos

Each video makes exactly one reproducible, measured claim, and every figure maps to a signed report in theEvidence Library. The benchmark suite is open source and reproducible:mingxin-kvcache-bench.

KV Cache Tiering on AMD MI308X: +29-40% LLM Throughput (Signed Benchmarks)

How NVMe-oF all-flash KV cache tiering lifts LLM inference throughput by 29-40% and cuts TTFT by 26-32% on an 8x AMD MI308X node (480B model, TP8). Every number carries a signed test-report citation.

English narration · 70 seconds · download MP4

Loading a 480B Model 6.2-9.3x Faster with All-Flash NVMe-oF (Measured)

Cold-starting large models is a storage problem. Measured on 8x AMD MI308X: all-flash NVMe-oF cuts 480B model load time by 6.2-9.3x versus the baseline path, which changes what elastic GPU scheduling can do.

English narration · 57 seconds · download MP4

Disclosure: narration and visuals in these videos are AI-generated. All performance figures come from signed third-party test reports and are reproduced without embellishment.