Getting 10 tok/s on Qwen 3.8 Flash Next: Why N-Gram Disk Paging Works Where MoE Fails
How I got Qwen 3.8 Flash Next running at 10 tokens/sec on an RTX 5060 Ti, and why paging n-gram tables from NVMe avoids the brutal disk bandwidth bottlenecks of MoE expert streaming.

