Lightweight PDF parser with layout, tables, formulas and bounding boxes
61 points by beatrizalmeidaf 6 hours ago | 7 comments

antonly 4 hours ago
I just tried this and the results look very off. I put in the MiniSat[1] Paper, and the result appears catastrophically wrong: https://imgur.com/a/i3RHwcQ

[1]: http://minisat.se/downloads/MiniSat.pdf

reply
phenomen 5 hours ago
I currently use https://github.com/firecrawl/anydoc in my PDF pipelines. In most cases it performs well. I'll test your lib to compare.
reply
coinfused 5 hours ago
Do you know if this one can also crop the PDF to a region of interest before converting to markdown?
reply
JaumeGar 4 hours ago
Really nice work, thanks for sharing. The table and formula extraction look great.
reply
archeantus 5 hours ago
Great work, thanks for sharing.
reply
beatrizalmeidaf 6 hours ago
I built this because extracting PDFs into Markdown/JSON often loses reading order, tables, formulas, figures, and their original locations.

The goal is a lightweight document extraction pipeline that preserves document structure and bounding boxes while exporting to Markdown, JSON, Excel and Word.

I'm also working on structure-aware semantic chunking for RAG, so retrieved chunks can retain their section, page and exact visual location in the PDF.

The project is open source and I'd love feedback on the architecture, extraction quality, and useful use cases.

reply
thatcherc 5 hours ago
This looks fantastic! The table and formula extraction features are especially interesting. My immediate question is: can this be integrated into Zotero? Most of the PDFs I read are research papers and extracting tables and formulas directly from my zotero collection would be super handy.
reply