Pure-C inference engine with zero dependencies for running very large mixture-of-experts models on commodity hardware by streaming experts from disk across unified memory and storage layers.