Our tools — open source
Reads documents other tools can't.
A local-first desktop app that turns PDFs, DOCX, PPTX and more into clean, AI-ready Markdown — with built-in OCR that recovers the text other tools return empty.
OCR built in
Reads scanned, image-only PDFs that other extractors return empty.
Local-first & private
Runs entirely on-device. No API key, nothing uploaded, nothing leaves your machine.
Clean, AI-ready Markdown
Headings, lists and tables survive intact. Read it, diff it, feed it to any model.
8+ formats, one drop
PDF · DOCX · PPTX · XLSX · EPUB · HTML & more — drag a folder, get Markdown.
vs. Microsoft MarkItDown
Built for the documents MarkItDown leaves behind.
Scanned PDFs come back empty
pdfminer only reads an embedded text layer.
Built-in OCR recovers the text
On-device — the scans others return blank.
OCR needs your own cloud key
Plugin route — cost, latency, pages leave the machine.
Local OCR, no key
No API key, no per-image bill, nothing uploaded.
Structure flattens to plain text
Heuristic extraction loses headings & tables.
Structure stays intact
Headings, lists & tables survive into the Markdown.
A Python library & CLI
pip, dependencies, Python 3.10+ to get started.
A local-first desktop app
Drag a folder, get Markdown. No setup.
Same MIT-licensed, local-by-default spirit — MDFlux just doesn't stop at the embedded text layer.