Reusable persistent KV-cache management layer for LLM inference, compatible with multiple engines and storage backends, helping engineers reduce time-to-first-token and improve throughput.