Maintains a fork of the llama.cpp local LLM inference engine, supporting low-bit model formats and multiple hardware backends for developers building local inference services.