An inference engine built on stock llama.cpp that runs oversized MoE models on memory-constrained phones by streaming only routed experts from flash without quality loss.