Fast LLM inference server for consumer-grade GPUs that uses speculative decoding and custom kernels to raise throughput while lowering memory use.