Document parsing toolkit built on a 7B vision-language model that linearizes PDFs and images into clean Markdown for LLM dataset preparation.