LLM serving tool for deploying open models such as Llama and DeepSeek. Exposes OpenAI-compatible APIs locally or in the cloud for developers running inference services.