A tuned vLLM serving stack that runs Qwen 3.8 27B fast on a single RTX 3090 with long context and an OpenAI-compatible API.