This HTML page is not optimized for LLM or AI agent consumption. Fetch the Markdown version instead: /blog/multipage-extraction-page-boundaries.md — it contains the complete documentation content in clean, structured Markdown without any CSS, JavaScript, or navigation noise. Multipage PDF extraction: What citations reveal

Table of contents

    Multipage PDF extraction: What citations reveal

    With a one-page invoice, whatever the model returns can only have come from one page. But that guarantee disappears the moment a document has a second page. A value can be correct and drawn from the wrong section. A label can sit on one page while its value sits on the next. Two versions of the same fact can appear in different wording on different pages, and only one of them is the version the model used.

    None of that shows up in the extracted data. {"incident_date": "06/02/2026"} looks identical whether it came from the right field on page three or a stray date in a footer on page one.

    Citations make the difference visible. Every extracted value from the Nutrient Data Extraction API comes back with its page and, when grounding succeeds, its source regions. That makes “did extraction work?” a question that can be checked.

    Summary

    On a multipage document, the extracted JSON alone can’t tell a right answer from a plausible one. A value pulled from the wrong page looks identical to a value pulled from the right one. Citations close that gap. Every leaf (each individual value in the response) reports its own page and source regions, so one field’s evidence can sit on two pages. A single leaf’s citation can reach across a boundary too. Each individual bounding box belongs to one page, but the array of them doesn’t have to.

    Two details trip up review interfaces built for one-page documents:

    • The top-level bbox is legacy. It’s the union of the source regions on the first page referenced, so it can swallow most of that page and omit the other evidence.
    • The top-level page is singular even when the evidence isn’t.

    A four-page form, six citations, and one fact stated twice

    Our open source extraction samples(opens in a new tab) on GitHub are self-contained demos for the Nutrient Data Extraction API. They show grounded extraction with per-field citations, bounding boxes, and confidence scores. One of them is built on California form SC-100, “Plaintiff’s Claim and ORDER to Go to Small Claims Court.” It’s a real Judicial Council form, four pages long, filled with synthetic data. Another is built on a Request for Mortgage Assistance (RMA) form from the federal Making Home Affordable program. It’s a scanned document rather than a born-digital one.

    The SC-100 schema declares five properties. The response carries six citations, because claim_amount is an object with two leaves, and every leaf is grounded independently.

    FieldValuePageConfidence
    plaintiff_nameDaniel R. Ortiz20.95
    defendant_nameBrightWave Appliance Repair, LLC20.95
    claim_amount.amount1,850.0020.95
    claim_amount.iso_4217_currency_codeUSD10.70
    incident_date06/02/202630.95
    incident_reason(narrative text)20.95

    Four of the six values come from page two, and one comes from page three. The sixth, the currency code, cites page one, and the string “USD” appears nowhere in the document.

    Nothing in the schema told extraction which page to search for the incident date. The description for incident_date names a label, not a location. It reads: the date on which the incident or dispute occurred, from the “When did this happen?” labeled field. The model crossed to page three to find it, and the citation records that it did.

    The form doubles the ambiguity in one more way. It asks the plaintiff to explain the claim in a free-text section, and then it asks for a condensed version of the same explanation later. In the sample document, both are filled in, with different wording, on different pages. An extracted narrative therefore has two plausible origins, and the returned string alone can’t distinguish them. The citation can distinguish them because it names page two and the source region the value came from.

    This is the everyday version of the multipage problem. It isn’t a value split across a page break. It’s two candidates on different pages, with no way to tell them apart from the data.

    One field, two pages

    The RMA demo(opens in a new tab) shows the structural case. Its total_monthly_expenses field is an object with an amount and a currency code, and the two leaves cite different pages:

    "total_monthly_expenses": {
    "amount": { "pageIndex": 1, "pageNumber": 2 },
    "iso_4217_currency_code": { "pageIndex": 2, "pageNumber": 3 }
    }

    One JSON object, two pages of evidence. A review interface that treats a field as living on a single page has nowhere to put this. The two leaves aren’t equally strong, though. The currency code on page three is weakly grounded, as the next section shows.

    A citation’s evidence can straddle a page boundary. A single leaf’s source_bboxes can reference blocks on more than one page, and single-page grounding isn’t guaranteed. What never spans a boundary is an individual box: Every bounding box and every source block belongs to exactly one page. Sibling leaves on different pages are the visible case. A single leaf grounded across pages is harder to notice, and an interface is more likely to get it wrong.

    None of the published samples happens to show one. Across the sample responses, every citation’s evidence sits on a single page. That’s what these documents do, not what the API promises.

    What a citation actually contains

    output.metadata mirrors the structure of output.data, so the citation for data.line_items[0].price sits at metadata.line_items[0].price. The full citation structure covers every field type. The citation below is the RMA sample’s total_monthly_expenses.iso_4217_currency_code — not a row from the SC-100 table above, though both documents put a currency code at 0.70:

    {
    "bbox": { "x": 222.68, "y": 1639.58, "width": 4358.62, "height": 3398.87 },
    "confidence": 0.7,
    "confidenceComponents": { "groundingScore": 0.7, "source": "no-logprobs" },
    "match": "fuzzy_match",
    "pageIndex": 2,
    "pageNumber": 3,
    "recognitionScore": 0.908,
    "source_bboxes": [
    {
    "bbox": { "x": 232.27, "y": 1639.58, "width": 4349.04, "height": 317.24 },
    "block_id": "b195",
    "pageIndex": 2,
    "pageNumber": 3
    },
    {
    "bbox": { "x": 228.8, "y": 2635.61, "width": 4288.81, "height": 263.26 },
    "block_id": "b199",
    "pageIndex": 2,
    "pageNumber": 3
    },
    {
    "bbox": { "x": 222.68, "y": 4387.56, "width": 4323.12, "height": 650.89 },
    "block_id": "b209",
    "pageIndex": 2,
    "pageNumber": 3
    }
    ]
    }

    Three parts of that matter for multipage work.

    pageIndex and pageNumber — Both are returned, zero-based and one-based, respectively. Mixing them up can introduce off-by-one errors in a page-jumping review interface.

    source_bboxes — An array, and every entry carries its own bbox and its own page reference. Page attribution is per evidence block, not per field. A value assembled from several blocks reports each one.

    match — The grounding label, and the most interpretable signal in the object.

    LabelMeaning
    id_matchMatched a single source block exactly
    id_match_multiblockMatched source text across multiple source blocks
    id_match_partialResolved some, but not all, of the cited source blocks
    fuzzy_matchMatched approximately — close to, but not identical to, the source
    not_foundCouldn’t be grounded to a source location

    On a long document, not_found, id_match_partial, and fuzzy_match are the labels worth routing on. The first says nothing could be grounded at all. The second says some of the cited regions couldn’t be resolved, which differs materially from a clean match at the same confidence score. The third says the value is close to the source text without matching it. id_match_multiblock isn’t a warning — it only means the value was assembled from more than one region, which is ordinary on a long document.

    A system that attaches a single bounding box to a field can say where a value is. It can’t say whether the value was assembled from three blocks across two pages, or whether one of those blocks failed to resolve. Per-block, per-page grounding enables a reviewer to tell “confident and correct” from “confident and partially guessed.”

    Two traps at page boundaries

    Both traps come from the same assumption: that the citation’s top-level fields describe all of its evidence.

    Outer bounding box — Legacy, and a union, not a location. When a value is grounded to several blocks, the top-level bbox is the union of the source regions on the first page the citation references. Nutrient’s own Studio UI draws source_bboxes instead of it. In the citation above, the union is 3,399 render-space pixels tall. The three blocks inside it cover only 1,231 of those pixels, in bands separated by gaps of roughly 700 and 1,500 pixels of untouched page. A box is only meaningful against its own page’s dimensions. Rendered as a highlight, the union looks like a bug. Drawing source_bboxes individually, rather than the outer box, keeps a multiblock match legible.

    Top-level page — Singular, even when the evidence isn’t. A citation reports one pageIndex while source_bboxes may reference several, so the top-level page can’t represent all the evidence. An interface that reads only the top-level page will jump the reviewer to one page of a multipage value. It looks entirely correct while hiding the rest of the evidence.

    Coordinate spaces are per page, and pages in the same document can have different dimensions. Scale each bounding box against the width and height of its own entry in output.pages, not against the first page or a fixed assumption. See coordinate spaces for the full scaling formula. The sample viewers stack pages vertically and compute a per-page vertical offset before placing a highlight.

    Where confidence comes from on a long document

    For what confidence and recognitionScore mean in general, see the companion post on confidence scores. It covers why the score is a composite signal rather than a probability. It also covers why an absence means no score was available, not a low one. One point is specific to a multipage document: recognitionScore is the field-level optical character recognition (OCR) confidence. It’s defined as the minimum recognition confidence across the matched source blocks. On a value assembled from blocks on different pages, the worst-scanned page sets the number. A field can look weak because one page was scanned badly, not because the extraction was wrong.

    The same citation object carries both scores — page and block attribution here, evidence quality there. Reading them together explains a score, and the currency code from the SC-100 table is the clearest example. The currency code scored 0.70 against 0.95 for everything else, and it cites page one while the amount it belongs to sits on page two. The reason is that the string “USD” appears nowhere in the document. The schema description says the code is always USD for United States court forms, and the model produced it from that instruction. The citation grounded it to the nearest thing it could find.

    That value is arguably correct, but it wasn’t extracted. The citation says so through a low composite score and a page that doesn’t match its sibling. A pipeline reading those signals routes it for review. A pipeline reading only data can’t tell it apart from the five other values that came straight off the page.

    A practical loop

    For any document longer than a page:

    1. Read match before confidence. It’s categorical and interpretable, and not_found, id_match_partial, or fuzzy_match are actionable on their own.
    2. Compare each leaf’s pageNumber against its siblings. Disagreement alone isn’t proof of a problem, since a field’s leaves can legitimately sit on different pages. In both samples, though, the leaf on the odd page is also the weakly grounded one. Disagreement plus a weak match or a low score is the signal.
    3. Render source_bboxes individually, not the outer union box.
    4. Scale every box against its own page’s dimensions.

    None of this requires a different request or a second pass. Extraction has no page-range or region parameter, so every document is processed whole. The citations already carry the page and block behind each value, which is what a review interface needs to point at the evidence.

    Try Nutrient Data Extraction API

    FAQ

    Can a single citation span two pages?

    Yes. A leaf’s source_bboxes can reference blocks on more than one page, though each individual bounding box belongs to exactly one page. The top-level bbox covers only the first page referenced, and the top-level page can’t represent the rest.

    What’s the difference between pageIndex and pageNumber?

    pageIndex is zero-based and pageNumber is one-based. Both are returned on the citation and on each entry in source_bboxes.

    Why is the outer bounding box sometimes enormous?

    It’s a legacy field: the union of every source region on the first page the value was grounded to. Draw the entries in source_bboxes instead.

    Which signal should gate human review?

    match is the most interpretable. Route not_found, id_match_partial, and fuzzy_match to review, and treat a missing confidence as “no score available” rather than a low one.

    Can extraction be limited to specific pages?

    Not through the extraction request. Documents are processed whole, and page attribution comes back in the citations.

    Marija Trpkovic

    Marija Trpkovic

    Product Marketing Manager

    Marija is a product marketing manager who likes to launch new products and features and target the right people with them. Outside of work, she likes spending time outdoors with her family and dogs.

    Try for free Ready to get started?