---
title: "Multipage PDF extraction: What citations reveal"
canonical_url: "https://www.nutrient.io/blog/multipage-extraction-page-boundaries/"
md_url: "https://www.nutrient.io/blog/multipage-extraction-page-boundaries.md"
last_updated: "2026-10-09T22:18:52.577Z"
description: "Multipage PDF extraction returns a page and bounding box for every value. See how citations track fields across page boundaries in a real four-page form."
---

With a one-page invoice, whatever the model returns can only have come from one page. But that guarantee disappears the moment a document has a second page. A value can be correct and drawn from the wrong section. A label can sit on one page while its value sits on the next. Two versions of the same fact can appear in different wording on different pages, and only one of them is the version the model used.

None of that shows up in the extracted data. `{"incident_date": "06/02/2026"}` looks identical whether it came from the right field on page three or a stray date in a footer on page one.

Citations make the difference visible. Every extracted value from the [Nutrient Data Extraction API](https://www.nutrient.io/api/data-extraction-api/) comes back with its page and, when grounding succeeds, its source regions. That makes “did extraction work?” a question that can be checked.

**TL;DR**

On a multipage document, the extracted JSON alone can’t tell a right answer from a plausible one. A value pulled from the wrong page looks identical to a value pulled from the right one. Citations close that gap. Every leaf (each individual value in the response) reports its own page and source regions, so one field’s evidence can sit on two pages. A single leaf’s citation can reach across a boundary too. Each individual bounding box belongs to one page, but the array of them doesn’t have to.

Two details trip up review interfaces built for one-page documents:

- The top-level `bbox` is legacy. It’s the union of the source regions on the _first_ page referenced, so it can swallow most of that page and omit the other evidence.

- The top-level page is singular even when the evidence isn’t.

## A four-page form, six citations, and one fact stated twice

Our [open source extraction samples](https://github.com/PSPDFKit/nutrient-extraction-samples/) on GitHub are self-contained demos for the Nutrient Data Extraction API. They show grounded extraction with per-field citations, bounding boxes, and confidence scores. One of them is built on California form SC-100, “Plaintiff’s Claim and ORDER to Go to Small Claims Court.” It’s a real Judicial Council form, four pages long, filled with synthetic data. Another is built on a Request for Mortgage Assistance (RMA) form from the federal Making Home Affordable program. It’s a scanned document rather than a born-digital one.

The SC-100 schema declares five properties. The response carries six citations, because `claim_amount` is an object with two leaves, and every leaf is grounded independently.

| Field                                 | Value                            | Page | Confidence |
| ------------------------------------- | -------------------------------- | ---- | ---------- |
| `plaintiff_name`                      | Daniel R. Ortiz                  | 2    | 0.95       |
| `defendant_name`                      | BrightWave Appliance Repair, LLC | 2    | 0.95       |
| `claim_amount.amount`                 | 1,850.00                         | 2    | 0.95       |
| `claim_amount.iso_4217_currency_code` | USD                              | 1    | 0.70       |
| `incident_date`                       | 06/02/2026                       | 3    | 0.95       |
| `incident_reason`                     | _(narrative text)_               | 2    | 0.95       |

Four of the six values come from page two, and one comes from page three. The sixth, the currency code, cites page one, and the string “USD” appears nowhere in the document.

Nothing in the schema told extraction which page to search for the incident date. The description for `incident_date` names a label, not a location. It reads: *the date on which the incident or dispute occurred, from the “When did this happen?” labeled field*. The model crossed to page three to find it, and the citation records that it did.

The form doubles the ambiguity in one more way. It asks the plaintiff to explain the claim in a free-text section, and then it asks for a condensed version of the same explanation later. In the sample document, both are filled in, with different wording, on different pages. An extracted narrative therefore has two plausible origins, and the returned string alone can’t distinguish them. The citation can distinguish them because it names page two and the source region the value came from.

This is the everyday version of the multipage problem. It isn’t a value split across a page break. It’s two candidates on different pages, with no way to tell them apart from the data.

## One field, two pages

The [RMA demo](https://github.com/PSPDFKit/nutrient-extraction-samples/tree/main/demos/rma_extraction) shows the structural case. Its `total_monthly_expenses` field is an object with an amount and a currency code, and the two leaves cite different pages:

```json

"total_monthly_expenses": {
  "amount": { "pageIndex": 1, "pageNumber": 2 },
  "iso_4217_currency_code": { "pageIndex": 2, "pageNumber": 3 }
}

```

One JSON object, two pages of evidence. A review interface that treats a field as living on a single page has nowhere to put this. The two leaves aren’t equally strong, though. The currency code on page three is weakly grounded, as the next section shows.

A citation’s evidence can straddle a page boundary. A single leaf’s `source_bboxes` can reference blocks on more than one page, and single-page grounding isn’t guaranteed. What never spans a boundary is an individual box: Every bounding box and every source block belongs to exactly one page. Sibling leaves on different pages are the visible case. A single leaf grounded across pages is harder to notice, and an interface is more likely to get it wrong.

None of the published samples happens to show one. Across the sample responses, every citation’s evidence sits on a single page. That’s what these documents do, not what the API promises.

## What a citation actually contains

`output.metadata` mirrors the structure of `output.data`, so the citation for `data.line_items[0].price` sits at `metadata.line_items[0].price`. The full [citation structure](https://www.nutrient.io/guides/dws-data-extraction/extract/citations-and-confidence.md) covers every field type. The citation below is the RMA sample’s `total_monthly_expenses.iso_4217_currency_code` — not a row from the SC-100 table above, though both documents put a currency code at 0.70:

```json

{
  "bbox": { "x": 222.68, "y": 1639.58, "width": 4358.62, "height": 3398.87 },
  "confidence": 0.7,
  "confidenceComponents": { "groundingScore": 0.7, "source": "no-logprobs" },
  "match": "fuzzy_match",
  "pageIndex": 2,
  "pageNumber": 3,
  "recognitionScore": 0.908,
  "source_bboxes": [
    {
      "bbox": { "x": 232.27, "y": 1639.58, "width": 4349.04, "height": 317.24 },
      "block_id": "b195",
      "pageIndex": 2,
      "pageNumber": 3
    },
    {
      "bbox": { "x": 228.8, "y": 2635.61, "width": 4288.81, "height": 263.26 },
      "block_id": "b199",
      "pageIndex": 2,
      "pageNumber": 3
    },
    {
      "bbox": { "x": 222.68, "y": 4387.56, "width": 4323.12, "height": 650.89 },
      "block_id": "b209",
      "pageIndex": 2,
      "pageNumber": 3
    }
  ]
}

```

Three parts of that matter for multipage work.

**`pageIndex` and `pageNumber`** — Both are returned, zero-based and one-based, respectively. Mixing them up can introduce off-by-one errors in a page-jumping review interface.

**`source_bboxes`** — An array, and every entry carries its own `bbox` and its own page reference. Page attribution is per evidence block, not per field. A value assembled from several blocks reports each one.

**`match`** — The grounding label, and the most interpretable signal in the object.

| Label                 | Meaning                                                            |
| --------------------- | ------------------------------------------------------------------ |
| `id_match`            | Matched a single source block exactly                              |
| `id_match_multiblock` | Matched source text across multiple source blocks                  |
| `id_match_partial`    | Resolved some, but not all, of the cited source blocks             |
| `fuzzy_match`         | Matched approximately — close to, but not identical to, the source |
| `not_found`           | Couldn’t be grounded to a source location                          |

On a long document, `not_found`, `id_match_partial`, and `fuzzy_match` are the labels worth routing on. The first says nothing could be grounded at all. The second says some of the cited regions couldn’t be resolved, which differs materially from a clean match at the same confidence score. The third says the value is close to the source text without matching it. `id_match_multiblock` isn’t a warning — it only means the value was assembled from more than one region, which is ordinary on a long document.

A system that attaches a single bounding box to a field can say where a value is. It can’t say whether the value was assembled from three blocks across two pages, or whether one of those blocks failed to resolve. Per-block, per-page grounding enables a reviewer to tell “confident and correct” from “confident and partially guessed.”

## Two traps at page boundaries

Both traps come from the same assumption: that the citation’s top-level fields describe all of its evidence.

**Outer bounding box** — Legacy, and a union, not a location. When a value is grounded to several blocks, the top-level `bbox` is the union of the source regions on the first page the citation references. Nutrient’s own Studio UI draws `source_bboxes` instead of it. In the citation above, the union is 3,399 render-space pixels tall. The three blocks inside it cover only 1,231 of those pixels, in bands separated by gaps of roughly 700 and 1,500 pixels of untouched page. A box is only meaningful against its own page’s dimensions. Rendered as a highlight, the union looks like a bug. Drawing `source_bboxes` individually, rather than the outer box, keeps a multiblock match legible.

**Top-level page** — Singular, even when the evidence isn’t. A citation reports one `pageIndex` while `source_bboxes` may reference several, so the top-level page can’t represent all the evidence. An interface that reads only the top-level page will jump the reviewer to one page of a multipage value. It looks entirely correct while hiding the rest of the evidence.

Coordinate spaces are per page, and pages in the same document can have different dimensions. Scale each bounding box against the `width` and `height` of its own entry in `output.pages`, not against the first page or a fixed assumption. See [coordinate spaces](https://www.nutrient.io/guides/dws-data-extraction/parsing/coordinate-spaces.md) for the full scaling formula. The sample viewers stack pages vertically and compute a per-page vertical offset before placing a highlight.

## Where confidence comes from on a long document

For what `confidence` and `recognitionScore` mean in general, see the [companion post on confidence scores](https://www.nutrient.io/blog/document-extraction-confidence-scores.md). It covers why the score is a composite signal rather than a probability. It also covers why an absence means no score was available, not a low one. One point is specific to a multipage document: `recognitionScore` is the field-level optical character recognition (OCR) confidence. It’s defined as the *minimum* recognition confidence across the matched source blocks. On a value assembled from blocks on different pages, the worst-scanned page sets the number. A field can look weak because one page was scanned badly, not because the extraction was wrong.

The same citation object carries both scores — page and block attribution here, evidence quality there. Reading them together explains a score, and the currency code from the SC-100 table is the clearest example. The currency code scored 0.70 against 0.95 for everything else, and it cites page one while the amount it belongs to sits on page two. The reason is that the string “USD” appears nowhere in the document. The schema description says the code is always USD for United States court forms, and the model produced it from that instruction. The citation grounded it to the nearest thing it could find.

That value is arguably correct, but it wasn’t extracted. The citation says so through a low composite score and a page that doesn’t match its sibling. A pipeline reading those signals routes it for review. A pipeline reading only `data` can’t tell it apart from the five other values that came straight off the page.

## A practical loop

For any document longer than a page:

1. Read `match` before `confidence`. It’s categorical and interpretable, and `not_found`, `id_match_partial`, or `fuzzy_match` are actionable on their own.

2. Compare each leaf’s `pageNumber` against its siblings. Disagreement alone isn’t proof of a problem, since a field’s leaves can legitimately sit on different pages. In both samples, though, the leaf on the odd page is also the weakly grounded one. Disagreement _plus_ a weak `match` or a low score is the signal.

3. Render `source_bboxes` individually, not the outer union box.

4. Scale every box against its own page’s dimensions.

None of this requires a different request or a second pass. Extraction has no page-range or region parameter, so every document is processed whole. The citations already carry the page and block behind each value, which is what a review interface needs to point at the evidence.

**Call to Action**

Try Nutrient Data Extraction API

[Learn More](https://www.nutrient.io/api/data-extraction-api/)

## FAQ

#### Can a single citation span two pages?

Yes. A leaf’s `source_bboxes` can reference blocks on more than one page, though each individual bounding box belongs to exactly one page. The top-level `bbox` covers only the first page referenced, and the top-level page can’t represent the rest.

#### What’s the difference between <code>pageIndex</code> and <code>pageNumber</code>?

`pageIndex` is zero-based and `pageNumber` is one-based. Both are returned on the citation and on each entry in `source_bboxes`.

#### Why is the outer bounding box sometimes enormous?

It’s a legacy field: the union of every source region on the first page the value was grounded to. Draw the entries in `source_bboxes` instead.

#### Which signal should gate human review?

`match` is the most interpretable. Route `not_found`, `id_match_partial`, and `fuzzy_match` to review, and treat a missing `confidence` as “no score available” rather than a low one.

#### Can extraction be limited to specific pages?

Not through the extraction request. Documents are processed whole, and page attribution comes back in the citations.

## Related reading

- [Citations and confidence](https://www.nutrient.io/guides/dws-data-extraction/extract/citations-and-confidence.md) — The full citation object, match labels, and coordinate space

- [Coordinate spaces](https://www.nutrient.io/guides/dws-data-extraction/parsing/coordinate-spaces.md) — Scaling formulas and overlay examples

- [Define a schema](https://www.nutrient.io/guides/dws-data-extraction/extract/define-a-schema.md) — Supported keywords and how field descriptions steer extraction

- [What should a document extraction confidence score actually tell you?](https://www.nutrient.io/blog/document-extraction-confidence-scores.md) — Why a source-grounded score behaves differently from a model-decisiveness score

- [PDF data extraction developer guide](https://www.nutrient.io/blog/pdf-data-extraction-developer-guide.md) — Parse modes, schema design, and how citations fit the wider pipeline

- [Nutrient extraction samples](https://github.com/PSPDFKit/nutrient-extraction-samples/) — The open source demos, including the SC-100 case
---

## Related pages

- [The business case for accessibility: Five ways it drives enterprise value](/blog/5-ways-accessibility-drives-enterprise-value.md)
- [Accessibility Untangled Why It Matters Guide](/blog/accessibility-untangled-why-it-matters-guide.md)
- [Advanced Techniques For React Native Ui Components](/blog/advanced-techniques-for-react-native-ui-components.md)
- [`vector_store` holds your indexed documents (see the multimodal RAG post](/blog/agentic-rag.md)
- [Ai Assistant Agent Governance](/blog/ai-assistant-agent-governance.md)
- [How to build an AI agent for contract redlining against a compliance playbook](/blog/ai-contract-redlining-compliance-playbook.md)
- [Ai Document Automation Extraction To Action](/blog/ai-document-automation-extraction-to-action.md)
- [Ai Document Workflows Ocr Compliance Heavy Teams](/blog/ai-document-workflows-ocr-compliance-heavy-teams.md)
- [Ai Legal Assistant Document Authoring](/blog/ai-legal-assistant-document-authoring.md)
- [Ai Schema Generator Document Extraction](/blog/ai-schema-generator-document-extraction.md)
- [Amazon Textract Alternatives](/blog/amazon-textract-alternatives.md)
- [Start (clears any prior buffer), navigate the document, then stop into a file.](/blog/android-faster-pdf-rendering.md)
- [Android Pdf Out Of Memory Handling](/blog/android-pdf-out-of-memory-handling.md)
- [Angular File Viewer Pdf Image Office Files](/blog/angular-file-viewer-pdf-image-office-files.md)
- [Approval Workflow Software](/blog/approval-workflow-software.md)
- [Approvals Matrix](/blog/approvals-matrix.md)
- [Apryse To Nutrient Migration](/blog/apryse-to-nutrient-migration.md)
- [Auto Tagging And Document Accessibility In Dotnet Sdk](/blog/auto-tagging-and-document-accessibility-in-dotnet-sdk.md)
- [Simple PII redaction.](/blog/automated-pii-removal.md)
- [Azure Document Intelligence Alternatives](/blog/azure-document-intelligence-alternatives.md)
- [Best Ai Document Workflow Platforms](/blog/best-ai-document-workflow-platforms.md)
- [Best Document Ai Platforms](/blog/best-document-ai-platforms.md)
- [Best Document Classification Platforms](/blog/best-document-classification-platforms.md)
- [Best document parser for RAG: LlamaParse vs. Unstructured vs. Reducto vs. Nutrient](/blog/best-document-parser-llamaparse-unstructured-reducto.md)
- [Best Document Parsing Apis](/blog/best-document-parsing-apis.md)
- [Best Document Viewers](/blog/best-document-viewers.md)
- [Best Llm Document Understanding Platforms](/blog/best-llm-document-understanding-platforms.md)
- [Best Multilingual Ocr Software](/blog/best-multilingual-ocr-software.md)
- [Best Pdf Parsers For Rag](/blog/best-pdf-parsers-for-rag.md)
- [Best Salesforce Document Generation Apps](/blog/best-salesforce-document-generation-apps.md)
- [Best Secure Document Collaboration Platforms](/blog/best-secure-document-collaboration-platforms.md)
- [Bpm Guide](/blog/bpm-guide.md)
- [Bpm Tools](/blog/bpm-tools.md)
- [Build Vs Buy Document Extraction](/blog/build-vs-buy-document-extraction.md)
- [Business Automation](/blog/business-automation.md)
- [Capex Vs Opex](/blog/capex-vs-opex.md)
- [The CEO’s AI playbook: Why decision architecture beats model selection](/blog/ceo-ai-playbook-decision-architecture.md)
- [1. Extract and chunk the PDF.](/blog/chat-with-pdf.md)
- [Complete Guide To Pdfjs](/blog/complete-guide-to-pdfjs.md)
- [Construction Document Data Extraction](/blog/construction-document-data-extraction.md)
- [Convert One Drive Files To Pdf In Sharepoint](/blog/convert-one-drive-files-to-pdf-in-sharepoint.md)
- [Create And Edit Pdfs In Flutter](/blog/create-and-edit-pdfs-in-flutter.md)
- [Create Pdfs With React](/blog/create-pdfs-with-react.md)
- [Creating A Document Scanner With Ocr In Python](/blog/creating-a-document-scanner-with-ocr-in-python.md)
- [Creating And Filling Pdf Forms Programmatically In Javascript](/blog/creating-and-filling-pdf-forms-programmatically-in-javascript.md)
- [The CTO’s AI playbook: Why accountability architecture beats orchestration](/blog/cto-ai-playbook-accountability-architecture.md)
- [Digital Signatures](/blog/digital-signatures.md)
- [Digital Workflow Automation](/blog/digital-workflow-automation.md)
- [Docling Alternatives](/blog/docling-alternatives.md)
- [Document Ai Vs Ocr](/blog/document-ai-vs-ocr.md)
- [Document Authoring Audit Trail](/blog/document-authoring-audit-trail.md)
- [Document Extraction Confidence Scores](/blog/document-extraction-confidence-scores.md)
- [Document Extraction For Underwriting](/blog/document-extraction-for-underwriting.md)
- [Document Viewer](/blog/document-viewer.md)
- [Document Watermarking](/blog/document-watermarking.md)
- [Electronic signature API: How to sign PDFs programmatically](/blog/electronic-signature-api.md)
- [Emerging threats: Your logging system may be an agentic threat vector](/blog/emerging-threats-your-logging-system.md)
- [Extend Alternatives](/blog/extend-alternatives.md)
- [Extract Patient Data On Premises](/blog/extract-patient-data-on-premises.md)
- [app.py](/blog/extract-text-from-pdf-using-python.md)
- [Fillable Pdf](/blog/fillable-pdf.md)
- [Google Document Ai Alternatives](/blog/google-document-ai-alternatives.md)
- [How To Add Digital Signature To Pdf Using React](/blog/how-to-add-digital-signature-to-pdf-using-react.md)
- [How To Build A Dotnet Maui Pdf Viewer](/blog/how-to-build-a-dotnet-maui-pdf-viewer.md)
- [How To Build A Flutter Pdf Viewer](/blog/how-to-build-a-flutter-pdf-viewer.md)
- [or](/blog/how-to-build-a-javascript-pdf-viewer-with-pdfjs.md)
- [or](/blog/how-to-build-a-javascript-pdf-viewer.md)
- [How To Build A Nextjs Pdf Viewer](/blog/how-to-build-a-nextjs-pdf-viewer.md)
- [How To Build A Powerpoint Viewer Using Javascript](/blog/how-to-build-a-powerpoint-viewer-using-javascript.md)
- [Using Yarn](/blog/how-to-build-a-react-excel-viewer.md)
- [How To Build A React Native Pdf Viewer](/blog/how-to-build-a-react-native-pdf-viewer.md)
- [How To Build A React Powerpoint Viewer](/blog/how-to-build-a-react-powerpoint-viewer.md)
- [How To Build A Reactjs File Viewer](/blog/how-to-build-a-reactjs-file-viewer.md)
- [or](/blog/how-to-build-a-reactjs-pdf-viewer-with-react-pdf.md)
- [or](/blog/how-to-build-a-reactjs-pdf-viewer.md)
- [How To Build A Reactjs Viewer With Pdfjs](/blog/how-to-build-a-reactjs-viewer-with-pdfjs.md)
- [How To Build A Vuejs Pdf Viewer With Pdfjs](/blog/how-to-build-a-vuejs-pdf-viewer-with-pdfjs.md)
- [How To Build A Vuejs Pdf Viewer](/blog/how-to-build-a-vuejs-pdf-viewer.md)
- [How To Build An Android Pdf Viewer](/blog/how-to-build-an-android-pdf-viewer.md)
- [How To Build An Angular Pdf Viewer With Ng2 Pdf Viewer](/blog/how-to-build-an-angular-pdf-viewer-with-ng2-pdf-viewer.md)
- [How To Build An Angular Pdf Viewer With Pdfjs](/blog/how-to-build-an-angular-pdf-viewer-with-pdfjs.md)
- [How To Convert Docx To Pdf Using Javascript](/blog/how-to-convert-docx-to-pdf-using-javascript.md)
- [How To Convert Docx To Pdf Using Python](/blog/how-to-convert-docx-to-pdf-using-python.md)
- [How To Convert Html To Pdf Using Html2pdf](/blog/how-to-convert-html-to-pdf-using-html2pdf.md)
- [How To Convert Html To Pdf Using React](/blog/how-to-convert-html-to-pdf-using-react.md)
- [How To Convert Html To Pdf Using Wkhtmltopdf And Csharp](/blog/how-to-convert-html-to-pdf-using-wkhtmltopdf-and-csharp.md)
- [or](/blog/how-to-convert-html-to-pdf-using-wkhtmltopdf-and-python.md)
- [How To Convert Html To Pptx](/blog/how-to-convert-html-to-pptx.md)
- [Quarterly report](/blog/how-to-convert-pdf-to-markdown-using-python.md)
- [How To Convert Word To Pdf In Nodejs](/blog/how-to-convert-word-to-pdf-in-nodejs.md)
- [or](/blog/how-to-create-a-react-js-signature-pad.md)
- [How To Create Pdfs With React To Pdf](/blog/how-to-create-pdfs-with-react-to-pdf.md)
- [How To Edit Pdfs Using Ios Pdf Library](/blog/how-to-edit-pdfs-using-ios-pdf-library.md)
- [How To Embed A Pdf Viewer In Your Website](/blog/how-to-embed-a-pdf-viewer-in-your-website.md)
- [How To Extract Tables From Pdf And Images](/blog/how-to-extract-tables-from-pdf-and-images.md)
- [How To Generate Pdf From Html With Nodejs](/blog/how-to-generate-pdf-from-html-with-nodejs.md)
- [How to generate PDFs in C# with Nutrient .NET SDK](/blog/how-to-generate-pdf-in-csharp.md)
- [base_url tells WeasyPrint where to resolve relative asset paths](/blog/how-to-generate-pdf-reports-from-html-in-python.md)
- [How To Merge Pdfs Using Javascript](/blog/how-to-merge-pdfs-using-javascript.md)
- [How To Ocr Pdfs In Linux](/blog/how-to-ocr-pdfs-in-linux.md)
- [How To Print Pdf In Csharp](/blog/how-to-print-pdf-in-csharp.md)
- [How To Programmatically Create And Fill Pdf Form In Angular](/blog/how-to-programmatically-create-and-fill-pdf-form-in-angular.md)
- [C# barcode scanner: How to scan and read 1D barcodes and QR codes](/blog/how-to-scan-barcodes-in-csharp.md)
- [Open an image.](/blog/how-to-use-tesseract-ocr-in-python.md)
- [From an HTML string.](/blog/html-in-pdf-format.md)
- [Html To Pdf In Javascript](/blog/html-to-pdf-in-javascript.md)
- [Intelligent Data Extraction](/blog/intelligent-data-extraction.md)
- [Invoice Approval Software](/blog/invoice-approval-software.md)
- [Javascript Document Editor](/blog/javascript-document-editor.md)
- [Javascript Pdf Editors](/blog/javascript-pdf-editors.md)
- [Javascript Pdf Libraries](/blog/javascript-pdf-libraries.md)
- [Landing Ai Alternatives](/blog/landing-ai-alternatives.md)
- [Langextract Vs Llamaindex Extraction Comparison](/blog/langextract-vs-llamaindex-extraction-comparison.md)
- [Linearized Pdf](/blog/linearized-pdf.md)
- [Uses OpenAI by default — set OPENAI_API_KEY.](/blog/llamaindex-vs-langchain-rag.md)
- [Llamaindex Vs Langchain Vs Haystack](/blog/llamaindex-vs-langchain-vs-haystack.md)
- [Llamaindex Workflows Vs Langgraph](/blog/llamaindex-workflows-vs-langgraph.md)
- [Llamaparse Alternatives](/blog/llamaparse-alternatives.md)
- [Low Code No Code Document Integrations](/blog/low-code-no-code-document-integrations.md)
- [Material Requisition](/blog/material-requisition.md)
- [or](/blog/merge-pdfs.md)
- [Migrating Observability Hyperdx To Dash0](/blog/migrating-observability-hyperdx-to-dash0.md)
- [Swift Package Manager](/blog/mobile-pdf-sdk.md)
- [`elements` come from your document parser — each has a type and content.](/blog/multimodal-rag.md)
- [Nutrient Flutter 6 Bindings Api](/blog/nutrient-flutter-6-bindings-api.md)
- [Nutrient Flutter Bindings Architecture](/blog/nutrient-flutter-bindings-architecture.md)
- [Nutrient Vs Conga Composer](/blog/nutrient-vs-conga-composer.md)
- [Ocr Api Comparison](/blog/ocr-api-comparison.md)
- [Online Document Viewer](/blog/online-document-viewer.md)
- [Open Pdf In Your Web App](/blog/open-pdf-in-your-web-app.md)
- [PDF accessibility for developers: Meeting WCAG 2.2, Section 508, and PDF/UA with an SDK](/blog/pdf-accessibility.md)
- [Extract data from PDF files: A developer guide to structured data from PDFs and scans](/blog/pdf-data-extraction-developer-guide.md)
- [Pdf Extraction Benchmark Opendataloader Bench](/blog/pdf-extraction-benchmark-opendataloader-bench.md)
- [Pdf Extraction Document Case Studies](/blog/pdf-extraction-document-case-studies.md)
- [Pdf Page Labels](/blog/pdf-page-labels.md)
- [Pdf Sdk Compliance Security Checklist](/blog/pdf-sdk-compliance-security-checklist.md)
- [Pdf Sdk Performance Benchmark](/blog/pdf-sdk-performance-benchmark.md)
- [Pdf Ua Compliance Guide](/blog/pdf-ua-compliance-guide.md)
- [Pdf Ua Validation](/blog/pdf-ua-validation.md)
- [Pdfjs Accessibility Structtree Printing](/blog/pdfjs-accessibility-structtree-printing.md)
- [Pdfjs Advanced Loading Streaming Workers](/blog/pdfjs-advanced-loading-streaming-workers.md)
- [Pdfjs Annotation Editor Layer](/blog/pdfjs-annotation-editor-layer.md)
- [Pdfjs Area Annotations Canvas Capture](/blog/pdfjs-area-annotations-canvas-capture.md)
- [Pdfjs Coordinate Systems Pdf To Screen](/blog/pdfjs-coordinate-systems-pdf-to-screen.md)
- [Pdfjs Document Outline Bookmarks Metadata](/blog/pdfjs-document-outline-bookmarks-metadata.md)
- [Pdfjs Eventbus Guide](/blog/pdfjs-eventbus-guide.md)
- [macOS](/blog/pdfjs-file-format-conversion-to-pdf.md)
- [macOS](/blog/pdfjs-generating-pdf-thumbnails-pdf2pic.md)
- [Pdfjs Limitations Commercial Upgrade](/blog/pdfjs-limitations-commercial-upgrade.md)
- [Pdfjs Native Annotation Layer Forms](/blog/pdfjs-native-annotation-layer-forms.md)
- [Pdfjs Navigation Zoom Rotation](/blog/pdfjs-navigation-zoom-rotation.md)
- [Pdfjs Pdf Page Manipulation Pdf Lib](/blog/pdfjs-pdf-page-manipulation-pdf-lib.md)
- [Pdfjs React Viewer Setup](/blog/pdfjs-react-viewer-setup.md)
- [Pdfjs Rendering Overlays React Portals](/blog/pdfjs-rendering-overlays-react-portals.md)
- [Pdfjs Server Side Text Extraction](/blog/pdfjs-server-side-text-extraction.md)
- [Pdfjs Sticky Note Annotations](/blog/pdfjs-sticky-note-annotations.md)
- [Pdfjs Text Highlight Annotations](/blog/pdfjs-text-highlight-annotations.md)
- [Pdfjs Text Search Pdffindcontroller](/blog/pdfjs-text-search-pdffindcontroller.md)
- [Pdfjs Thumbnail Sidebar](/blog/pdfjs-thumbnail-sidebar.md)
- [People Process Tools](/blog/people-process-tools.md)
- [Process Flows](/blog/process-flows.md)
- [React Native Pdf Annotation](/blog/react-native-pdf-annotation.md)
- [React Pdf Annotation Layer Forms](/blog/react-pdf-annotation-layer-forms.md)
- [React Pdf Custom Rendering Hooks](/blog/react-pdf-custom-rendering-hooks.md)
- [Using Yarn](/blog/react-pdf-editor.md)
- [React Pdf Loading States Errors Passwords](/blog/react-pdf-loading-states-errors-passwords.md)
- [React Pdf Non Latin Fonts Special Pdfs](/blog/react-pdf-non-latin-fonts-special-pdfs.md)
- [React Pdf Outline Table Of Contents](/blog/react-pdf-outline-table-of-contents.md)
- [React Pdf Performance Optimization](/blog/react-pdf-performance-optimization.md)
- [React Pdf Setup Basic Rendering](/blog/react-pdf-setup-basic-rendering.md)
- [React Pdf Text Layer Custom Renderer](/blog/react-pdf-text-layer-custom-renderer.md)
- [React Pdf Thumbnails Page Navigation](/blog/react-pdf-thumbnails-page-navigation.md)
- [Reducto Alternatives](/blog/reducto-alternatives.md)
- [Requisition System](/blog/requisition-system.md)
- [labels.py](/blog/route-documents-automatically-classify-api.md)
- [or](/blog/sample-blog-updated.md)
- [Sdk Product Updates Q2 2026](/blog/sdk-product-updates-q2-2026.md)
- [System Of Record Vs Source Of Truth](/blog/system-of-record-vs-source-of-truth.md)
- [Add DWS MCP Server to your Claude Code project.](/blog/teaching-llms-to-read-pdfs.md)
- [Open an image file.](/blog/tesseract-python-guide.md)
- [The Six Best Pdf Generator Apis](/blog/the-six-best-pdf-generator-apis.md)
- [Define the HTML part of the document.](/blog/top-10-ways-to-generate-pdfs-in-python.md)
- [Top 5 Javascript Pdf Viewers](/blog/top-5-javascript-pdf-viewers.md)
- [or](/blog/top-js-pdf-libraries.md)
- [Top Ten Ways To Convert Html To Pdf](/blog/top-ten-ways-to-convert-html-to-pdf.md)
- [Unstructured Alternatives](/blog/unstructured-alternatives.md)
- [Using Javascript In Pdf Form Fields](/blog/using-javascript-in-pdf-form-fields.md)
- [Vector Pdf](/blog/vector-pdf.md)
- [Wcag2 Accessibility Requirements Documents](/blog/wcag2-accessibility-requirements-documents.md)
- [Web Sdk Is Now Headless](/blog/web-sdk-is-now-headless.md)
- [What Are Annotations](/blog/what-are-annotations.md)
- [What Is A Vpat](/blog/what-is-a-vpat.md)
- [What Is Business Logic](/blog/what-is-business-logic.md)
- [What Is Document Processing](/blog/what-is-document-processing.md)
- [What Is Intelligent Document Processing](/blog/what-is-intelligent-document-processing.md)
- [What Is Ocr Invoice Processing](/blog/what-is-ocr-invoice-processing.md)
- [What Is Pdf Ua](/blog/what-is-pdf-ua.md)
- [Why Pdfium Is A Trusted Platform For Pdf Rendering](/blog/why-pdfium-is-a-trusted-platform-for-pdf-rendering.md)
- [Why Your Ai Agent Hallucinates Pdf Table Data](/blog/why-your-ai-agent-hallucinates-pdf-table-data.md)

