Python API and C++ runtime for efficient LLM inference optimization with specialized kernels that accelerate large model execution on NVIDIA GPUs.