Open-source inference serving system for deploying models from multiple AI frameworks across cloud and edge, supporting concurrent batching and streaming for ML platform engineers.