A reproducible guide and evaluation wiki documenting how to run large language models such as Qwen, Kimi, and GLM variants on RTX 6000 Pro PCIe GPUs without NVLink for local inference practitioners.