llama.cpp fork focused on local LLM inference efficiency, using KVarN, KV-cache precision tailoring, and low-bit quantization to support longer context with better precision in the same VRAM.