Comprehensive document-parsing benchmark with 1,651 pages of annotated real-world data plus end-to-end evaluation code for researchers evaluating layout analysis, OCR, and structured extraction systems.