· via Hacker News – Front Page (native)
papero: CPU-only PDF parser recovers tables, formulas and layout without ML models
An open-source parser called papero rebuilds PDF reading order, tables, formulas and per-block bounding boxes using geometry alone — no ML models, CPU only — in Python, the browser or as a Docker API.
Geometry-based PDF extraction without a GPU
An open-source project called papero surfaced on Hacker News's front page with an appealing pitch for anyone building document pipelines: reconstruct a PDF's structure — reading order, tables, formulas, figures and the position of every block — using geometry alone on a laptop CPU, with no ML models and no GPU. The repository (beatrizalmeidaf/papero-pdf-text-extractor) is MIT-licensed and can run as a Python library, a CLI, a Docker-hosted REST API, or entirely inside the browser.
According to the project's README, the problem being solved is not text extraction but structure recovery. Getting characters out of a PDF is easy; knowing which column to read first, which lines form a table, and where a formula sits is what makes output usable for RAG, search and LLM ingestion.
How it works
Two engines process the file in parallel. A layout engine built on PDFium reads every glyph with its position, font and size, plus rules and images, then rebuilds columns, tables, formulas, lists and figures using a column-aware XY-cut algorithm. Apache Tika contributes metadata, tagged-PDF heading detection, OCR via Tesseract, and parsing of non-PDF formats such as DOCX, PPTX, XLSX, EPUB and HTML; a tika=False option runs the layout engine alone without Java.
The browser app is a JavaScript port of the same algorithm running on pdf.js, and the project's CI checks block by block that the Python and browser engines agree — a useful guard against silent divergence between the two implementations.
Several edge cases are handled explicitly: accented characters drawn as separate glyphs in LaTeX-produced PDFs are reassembled, invisible white text used by form generators is dropped, and scanned pages go through OCR (Tesseract ships in the Docker image, defaulting to Portuguese and English).
Outputs and interfaces
Every block is typed — heading, paragraph, list item, table, figure, formula, caption, code — and carries a bounding box in points, while headers, footers and page numbers are kept separate from the body text. Formulas come back as approximate LaTeX for fractions, roots, exponents and indices, plus a cropped PNG as a fallback; figures are cropped to PNG with captions, axis labels and legends kept together; ruled, borderless and booktabs-style tables come back as rows and columns, exportable to CSV or Excel.
Developers install with pip (the package is named papero-extract and imports as pdf_text_api) and get Python bindings, a pdf-text-api CLI, and a docker compose stack exposing a POST /v1/extract endpoint plus a browser app on port 8000. The browser version processes files locally, so sensitive PDFs never leave the machine, and it also covers Word and Excel export, which is not yet available through the Python path.
Benchmarks and stated limits
The project's own benchmark, run on 54 dense arXiv papers with multiple columns, formulas, tables and figures, reports zero failures and a median of 39 ms per page on a laptop CPU with no GPU. Its comparison table positions papero against PyMuPDF, pdfplumber, pypdf, Docling and Marker: it is the only tool in that set that runs entirely in the browser, and unlike Docling and Marker it needs no PyTorch; PyMuPDF is AGPL and Marker is GPL, both awkward licenses for embedding in commercial products.
The limitations section is unusually specific. Mathematics is rebuilt from glyphs and strokes, so matrices, aligned equation systems and nested constructs come out linearized, with the cropped image always attached. Borderless tables with very narrow column gaps can be misread as plain text. Word export pins each PDF page to its own page and can spill a few lines onto an extra one. The README also concedes that ML-based tools still win on very irregular layouts and complex mathematics.
Why it matters
Document ingestion for RAG and LLM pipelines has drifted toward heavyweight ML stacks that assume a GPU and large model downloads. papero is a counterpoint: fast enough for interactive use, CPU-friendly, permissively licensed, and privacy-preserving when run in the browser. Per-block bounding boxes matter beyond layout, too — they let retrieval systems cite the exact spot on a page where a passage appeared, a small but important building block for grounded answers. For developers processing mostly well-formed documents such as papers, reports and contracts rather than battered scans, a geometry-first tool with a clear list of what it cannot do is often the more dependable choice.
- #open-source
- #document-processing
- #rag
- #developer-tools