Lightweight LLM inference engine implemented from scratch in about 1200 lines of Python, compatible with vLLM interfaces and offering prefix caching, tensor parallelism, and fast offline inference.