Evaluation framework and open benchmark registry for large language models and systems, allowing researchers to create custom tests and run standardized assessments of model capabilities.