Colibrì vs. FreeToken: Can Consumer Hardware Really Run Frontier MoE Models?
An empirical comparison of Colibri vs FreeToken for edge MoE serving, breaking down RAM limits, NVMe Direct DMA, the WSL2 I/O tax, and real tokens/sec.
An empirical comparison of Colibri vs FreeToken for edge MoE serving, breaking down RAM limits, NVMe Direct DMA, the WSL2 I/O tax, and real tokens/sec.
Empirical benchmarks comparing Windows Dev Drive (ReFS) vs NTFS for local LLMs across cold TTFT, inference jitter, Colibri MoE streaming, and n-gram paging.
How I got Qwen 3.8 Flash Next running at 10 tokens/sec on an RTX 5060 Ti, and why paging n-gram tables from NVMe avoids the brutal disk bandwidth bottlenecks of MoE expert streaming.
How to securely publish local LM Studio models to the internet using a Dockerized LiteLLM proxy and Cloudflare Tunnels.
A hands-on setup guide for running Clawbot against local LM Studio models and securely exposing the web interface using Cloudflare Tunnels.