PDF extraction API for developers and automation teams

Use the PDF Extractor API to automate image, table, text, formula, caption, and OCR extraction from PDFs with API tokens and structured outputs.

Automate PDF element extraction

Submit PDFs over HTTPS, choose categories and OCR settings, and retrieve extracted assets for your product, data pipeline, or AI workflow.

Use the same categories as the web app

API jobs can extract images, tables, text, titles, section titles, captions, footnotes, formulas, entries, headers, and footers.

Connect AI agents through MCP

Use the MCP server or direct REST calls to give assistants access to PDF images, tables, OCR text, and layout-aware context.

Configure each PDF extraction request

Select page ranges, output image format and resolution, extraction categories, confidence settings, and page-level or element-level OCR so the API returns only the assets the workflow needs.

Use predictable result files and metadata

Outputs carry page, category, and reading-order information in their filenames. Job metadata and result URLs make it practical to move files into storage, review queues, or AI-processing pipelines.

Prototype in the web app before automating

Teams can test category and OCR choices visually, then apply the same document concepts through API tokens or the PDF Extractor MCP server for repeatable application and agent workflows.