AI工具

vllm-2080ti-definitive

weicj/vllm-2080ti-definitive

The definitive vLLM runtime for dual RTX 2080 Ti 22GB + NVLink, delivering Qwen 27B local inference with maximum 100+ tok/s single-request decode with support of FP8 weight