PDF OCR for scanned documents and extracted elements

Run OCR on scanned PDFs and extracted elements. Export page text as Markdown, JSON, or plain text while also extracting images, tables, captions, and formulas.

Make scanned PDFs searchable

Use full-page OCR to create readable transcriptions from scans, legacy documents, manuals, contracts, and research papers.

Choose page OCR or element OCR

Run OCR on the entire page when you need reading order, or on detected elements when text should stay attached to tables, captions, and figures.

Export text in workflow-friendly formats

Download OCR output as Markdown, JSON, or plain text alongside extracted images, tables, formulas, and captions.

Use OCR without giving up page structure

OCR runs alongside layout detection, so a scanned page can produce a readable transcription plus separate image, table, formula, caption, title, header, and footer outputs from the same run.

Choose between standard and advanced OCR

Standard OCR is a practical choice for straightforward text. Advanced OCR is designed for layouts that benefit from Markdown or JSON, including headings, tables, multi-column pages, and figure markers.

Run scanned PDF OCR in the browser, API, or desktop app

Use the web interface for review, the REST API for recurring OCR pipelines, or the offline desktop app when source documents must remain on the local machine.