A fast Apple Silicon inference engine that runs local LLMs with exact speculative decoding behind an OpenAI-compatible API.