← Back to brief
ResearchOfficialPreprintarXiv Computation and Language

HPD-Parsing: Hierarchical Parallel Document Parsing Achieves 2.6x Throughput

Researchers introduce HPD-Parsing, a hierarchical parallel decoding paradigm for vision-language model (VLM)-based document parsing. By replacing full-page autoregressive generation with a main layout branch and concurrent block-level decoders, HPD-Parsing achieves 4,752 tokens per second—2.62 times the throughput of the fastest existing document parsing model—while maintaining competitive accuracy.

Why it matters: This work demonstrates a significant advance in document parsing efficiency, potentially enabling much faster processing of complex documents.

Full story at: arXiv Computation and Language