中文
Agent infrastructure

qwen38-27b-rtx3090

syv-ai/qwen38-27b-rtx3090

A tuned vLLM serving stack that runs Qwen 3.8 27B fast on a single RTX 3090 with long context and an OpenAI-compatible API.

Topics

  • kv-cache
  • llm-inference
  • local-llm
  • quantization
  • qwen
  • qwen3
  • rtx-3090
  • speculative-decoding