A library providing environments and infrastructure for evaluating and improving models and agents at scale. Supports reproducible benchmarks and reinforcement-training integrations across multiple frameworks for AI researchers.