PDF extraction API for developers and automation teams
Use the PDF Extractor API to automate image, table, text, formula, caption, and OCR extraction from PDFs with API tokens and structured outputs.
Automate PDF element extraction
Submit PDFs over HTTPS, choose categories and OCR settings, and retrieve extracted assets for your product, data pipeline, or AI workflow.
Use the same categories as the web app
API jobs can extract images, tables, text, titles, section titles, captions, footnotes, formulas, entries, headers, and footers.
Connect AI agents through MCP
Use the MCP server or direct REST calls to give assistants access to PDF images, tables, OCR text, and layout-aware context.
Configure each PDF extraction request
Select page ranges, output image format and resolution, extraction categories, confidence settings, and page-level or element-level OCR so the API returns only the assets the workflow needs.
Use predictable result files and metadata
Outputs carry page, category, and reading-order information in their filenames. Job metadata and result URLs make it practical to move files into storage, review queues, or AI-processing pipelines.
Prototype in the web app before automating
Teams can test category and OCR choices visually, then apply the same document concepts through API tokens or the PDF Extractor MCP server for repeatable application and agent workflows.