Low-memory inference library using layer-wise offloading to run 70B to 671B large language models for inference on a single GPU with limited VRAM, aimed at researchers and individual developers.