Provides a benchmark for evaluating and improving AI agents on legal work, combining realistic task datasets with execution harnesses and scoring scaffolds for professional workflows.