Delivers GB-per-second language model tokenization compatible with popular tokenizers, designed for high-throughput preprocessing and inference pipelines for researchers and engineers needing fast text encoding at scale.