Extract tables from PDF with a table extraction API

Use Nutrient DWS as a table extraction API to convert PDF tables to Excel and JSON. Extract tables from PDF documents when your workflow needs line items, statements, reports, and other tabular content as structured output instead of page images or manual data entry. CSV and XML output are in development — join the waitlist on their pages.

EXTRACT FROM
EXTRACT TO


How to extract tables from PDF

Extracting tables from a PDF takes one REST request to the Processor API. Send the PDF to the API endpoint, set the output type — xlsx for Excel or json-content for structured JSON — and receive the converted tables in the response. No Microsoft Office license or local installation is needed, so PDF table extraction runs inside headless and containerized pipelines. Ready-to-run examples ship for JavaScript, Python, Java, C#, PHP, and curl, so a first request takes minutes rather than an integration project.

PDF table-to-Excel conversion

Set the output type to xlsx to convert PDF tables into a Microsoft Excel workbook with rows and columns preserved, ready for spreadsheets, reporting, and reconciliation. The result is an editable workbook, not a flat text dump, so figures stay addressable in formulas and pivot tables.

PDF table to JSON for ETL

Set the output type to json-content to return PDF tables and text as structured JSON, with plain text and structured text in the same response. Structured text carries characters, lines, words, and paragraphs with their page positions, so ETL jobs, search indexes, and downstream services can consume it without a parsing layer in between.

Scanned tables and complex layouts

Scanned and image-only PDFs have no text layer, so an OCR action in the same request recognizes the characters before the tables are converted. When the table structure itself matters, such as cell boundaries, reading order, or a confidence score per value, the Data Extraction API returns typed table elements with bounding boxes and confidence scores, which conversion output doesn’t provide.

Extract tables from PDF in Python

Every example on this page includes a Python tab, so table extraction drops into a Python service using the standard requests library, with no SDK to install. Teams that need table extraction to run on their own infrastructure rather than the cloud API can use the Python SDK for PDF table extraction instead.

Data Extraction API

Extract specific fields from any document

Define a schema and get back typed JSON with confidence scores and source context for every value.

Start free:

    • 5,000 credits per month
    • No credit card required
    • Schema generator in Studio

Security is our top priority

SOC 2 Type 2 audited

Nutrient’s infrastructure is SOC 2 Type 2 audited and GDPR-compliant. See our privacy policy and security documentation for details on data handling.

HTTPS encryption

All communication between your application and Nutrient is done via HTTPS to ensure your data is encrypted when it’s sent to us.

Safe payment processing

All payments are handled by Paddle. Nutrient DWS Processor API never has direct access to any of your payment data.

Frequently asked questions

What is the PDF table extraction API?

Nutrient DWS Processor API extracts tabular data from PDF documents and returns it as structured output — Excel (XLSX) or JSON — so line items, statements, and reports become machine-readable instead of flat page images.

Which output formats can I extract PDF tables to?

Extracted tables convert to Excel (XLSX) and JSON today, choosing the format that fits your analytics, reporting, ETL, or downstream system. CSV and XML output are in development. If you’re interested, join the waitlist on the PDF-to-CSV and PDF-to-XML pages.

How do I extract data from a PDF into a table?

Send the PDF to the Processor API over REST and request the structured output you need. Ready-to-use examples are available in JavaScript, Python, Java, PHP, and HTTP (curl) to get a first request running quickly.

Can I convert PDF tables to Excel?

Yes. The PDF-to-Excel output returns an XLSX file with rows and columns preserved, ready for spreadsheets, reporting, and further processing.

Does table extraction work on scanned PDFs?

Scanned documents can be processed by pairing extraction with the Processor API’s OCR step, so text in image-based pages is recognized before tables are extracted.

Do I need Microsoft Office or any local installation?

No. Extraction runs server-side through the cloud API with no Microsoft Office license or local install required, so it fits headless and containerized workflows.

Is there a free tier, and how is it priced?

You can start free with Nutrient DWS credits and test with Postman before moving to credit-based Processor API pricing. See the Processor API pricing page for current credit details.

Ready to try it?

Create an account to get your DWS Processor API key and start making API calls.