Generator producing original ARC-AGI-1-style reasoning tasks matched to the public evaluation distribution, helping researchers test genuine generalization of frontier models.