Delivers a local inference engine for DeepSeek 4 Flash and PRO models on consumer hardware, supporting Metal, CUDA, and ROCm backends with multi-GPU parallel serving for developers and researchers.