This HTML page is not optimized for LLM or AI agent consumption. Fetch the Markdown version instead: /guides/dws-data-extraction/extract/production-structured-extraction.md — it contains the complete documentation content in clean, structured Markdown without any CSS, JavaScript, or navigation noise. Production structured extraction

Use the Nutrient DWS Data Extraction API extract endpoint to return structured JSON from documents in production workflows. This guide covers schema validation, citations, confidence gates, retry handling, and reproducible logs.

Use /extraction/extract when you know the target fields before processing. For whole-document Markdown or spatial elements, use the parse endpoint instead.

Production request template

Start with an explicit API version, the stable engine, a strict schema, citations, and a stored run:

Terminal window
curl -X POST https://api.nutrient.io/extraction/extract \
-H "Authorization: Bearer your_api_key_goes_here" \
-H "x-nutrient-api-version: 2026-05-25" \
-H "x-nutrient-engine-version: stable" \
-H "Content-Type: application/json" \
-d '{
"url": "https://storage.example.com/invoice.pdf",
"schema": {
"type": "object",
"properties": {
"invoice_number": {
"type": "string",
"description": "Invoice identifier exactly as printed on the invoice."
},
"invoice_date": {
"type": "string",
"format": "date",
"description": "Invoice issue date in ISO date format."
},
"total_amount": {
"type": "number",
"description": "Final invoice total including tax."
},
"currency": {
"type": "string",
"enum": ["USD", "EUR", "GBP"],
"description": "Invoice currency as an ISO 4217 code."
}
},
"required": ["invoice_number", "total_amount", "currency"]
},
"instructions": "Extract values exactly as shown. Do not infer missing values.",
"parseConfig": {
"mode": "understand",
"options": { "language": "auto" }
},
"options": {
"strict": true,
"includeCitations": true,
"multimodal": false
},
"storeRun": true
}'

This configuration uses the following settings:

SettingPurpose
x-nutrient-api-version: 2026-05-25Pins the public API contract for the request.
x-nutrient-engine-version: stableUses the stable production engine instead of the nightly preview engine.
schemaDefines the exact JSON shape to return. The root schema is an object.
options.strict: trueConstrains structured output to the supplied schema.
options.includeCitations: trueReturns per-field grounding metadata in output.metadata. Citations are enabled by default, but this setting makes the contract visible.
parseConfig.mode: "understand"Uses the recommended mode for most forms, invoices, receipts, tables, and key-value extraction workflows.
parseConfig.options.language: "auto"Explicitly enables automatic language detection. When you omit language, optical character recognition (OCR) also detects the language automatically.
storeRun: trueStores the run when run storage is available. The response returns a top-level runId for audit and debugging workflows.

Pin versions for reproducibility

Send both version headers on every production request:

x-nutrient-api-version: 2026-05-25
x-nutrient-engine-version: stable

Use stable for production workflows. Use nightly only when you test preview engine behavior before a production rollout.

Log the version headers with each extraction result, because they help you reproduce a run and compare behavior after API or engine changes.

Define a closed JSON schema

The extract endpoint returns data that matches the JSON Schema you provide. Use an inline schema with a root object:

{
"type": "object",
"properties": {
"invoice_number": {
"type": "string",
"description": "Invoice identifier exactly as printed on the invoice."
},
"invoice_date": {
"type": "string",
"format": "date",
"description": "Invoice issue date in ISO date format."
},
"total_amount": {
"type": "number",
"description": "Final invoice total including tax."
},
"currency": {
"type": "string",
"enum": ["USD", "EUR", "GBP"]
}
},
"required": ["invoice_number", "total_amount", "currency"]
}

The supported schema keywords include type, properties, required, items, description, string enum, and format: "date". Don’t use unsupported JSON Schema features, such as $ref, $defs, composition keywords, conditionals, numeric ranges, or string formats other than date.

For the full list of supported keywords and limits, refer to the define a schema guide.

Use strict output and application validation

Set options.strict to true for production schema extraction:

{
"options": {
"strict": true,
"includeCitations": true
}
}

strict: true constrains the extraction model to emit JSON that conforms to the supplied schema. It improves generation behavior, but it doesn’t replace validation in your application.

Validate output.data with the same schema before you use the result downstream, and then apply business rules that aren’t part of the supported schema subset. These rules can include numeric ranges, identifier formats, currency rules, or cross-field checks:

import Ajv from "ajv";
const ajv = new Ajv();
const validate = ajv.compile(schema);
if (!validate(result.output.data)) {
throw new Error(`Schema validation failed: ${ajv.errorsText(validate.errors)}`);
}

Choose the parse mode

The extract endpoint runs a parse stage before schema extraction, and you configure that stage with parseConfig.mode.

ModeUse when
structureDocuments have clean, predictable layouts, and lower cost or lower latency matters most.
understandMost invoices, forms, receipts, tables, and key-value extraction workflows.
agenticDocuments require complex visual reasoning, such as degraded scans, dense layouts, or diagrams.

Start with understand for most structured extraction. Move to structure after you test a representative sample and confirm the output quality. Move to agentic when understand misses visual context or degraded content.

Set language to "auto" when you want automatic OCR language detection:

{
"parseConfig": {
"mode": "understand",
"options": { "language": "auto" }
}
}

When you omit language, OCR detects the language automatically. For known languages, use a language name, an ISO 639-2 code, an array, or a +-joined string. For details, refer to the parse configuration and supported languages guides.

Enable multimodal extraction only when needed

Set options.multimodal to true when the extraction depends on visual layout or page images:

{
"options": {
"multimodal": true
}
}

Multimodal extraction sends rendered page images to the extraction model with parsed text. Use it for visually complex layouts, image-heavy documents, difficult forms, or fields whose meaning depends on spatial layout.

Multimodal extraction increases cost and latency. Keep it disabled for text-centric documents when parsed text and layout structure provide enough context.

Gate results with citations and confidence

With options.includeCitations: true, output.metadata mirrors output.data. Each scalar field can include a citation with a match label, page reference, bounding box, and confidence score:

{
"output": {
"data": {
"invoice_number": "INV-2024-0042",
"total_amount": 1547.5
},
"metadata": {
"invoice_number": {
"match": "id_match",
"confidence": 0.93,
"pageNumber": 1,
"bbox": { "x": 878, "y": 268, "width": 82, "height": 25 }
},
"total_amount": {
"match": "fuzzy_match",
"confidence": 0.76,
"pageNumber": 1,
"bbox": { "x": 930, "y": 1200, "width": 96, "height": 28 }
}
}
}
}

Use both match and confidence in review logic:

  • Route fuzzy_match and not_found fields to review.
  • Treat missing citations for required fields as review or failure.
  • Set per-field thresholds. Identifiers, totals, and regulated fields usually need stricter thresholds than descriptive text.

Confidence values are relative signals, not calibrated probabilities. Calibrate thresholds on your own labeled sample set for each document type and field.

Use the following policy as a starting point for calibration:

ConfidenceExample action
>= 0.90Accept automatically.
0.70 to 0.89Accept with warning or route to review.
< 0.70Reject or route to manual review.

Implement the policy in application code:

const REQUIRED_CONFIDENCE = {
invoice_number: 0.9,
total_amount: 0.9,
currency: 0.85,
};
function assertConfidence(result) {
const metadata = result.output.metadata ?? {};
for (const [field, minConfidence] of Object.entries(REQUIRED_CONFIDENCE)) {
const citation = metadata[field];
if (!citation) {
throw new Error(`Missing citation for ${field}`);
}
if (citation.match === "fuzzy_match" || citation.match === "not_found") {
throw new Error(`Review required for ${field}: ${citation.match}`);
}
if (citation.confidence != null && citation.confidence < minConfidence) {
throw new Error(`Low confidence for ${field}`);
}
}
}

For nested objects and arrays, traverse output.data and output.metadata together. For traversal examples, refer to the citations and confidence guide.

Handle errors explicitly

Branch on the HTTP status code, and retry only transient errors.

StatusMeaningHandling
400Bad request, invalid schema, invalid modeDon’t retry. Fix the request.
401Invalid or missing API keyDon’t retry. Fix credentials.
402Insufficient creditsDon’t retry automatically.
408TimeoutRetry with exponential backoff.
413File too largeSplit, compress, or reject.
422Remote URL rejected or unavailableFix the URL or input source.
429Rate limitedRetry with exponential backoff.
500Internal processing errorRetry and keep the requestId.
503Backend unavailable or overloadedRetry with exponential backoff.

Use bounded retries with jitter for 408, 429, 500, and 503. Keep requestId in logs for support and debugging.

For error response shapes and troubleshooting details, refer to the error handling guide.

Log audit and reproducibility fields

Store enough information to reproduce and debug each extraction, but avoid logging sensitive document content unless you need it.

Log these fields for every production extraction:

  • requestId.
  • runId, when storeRun: true returns it.
  • x-nutrient-api-version.
  • x-nutrient-engine-version.
  • Parse mode and parse options.
  • Extract options, including strict, includeCitations, and multimodal.
  • Schema version and schema hash.
  • Instruction version and instruction hash.
  • Input document hash.
  • Application version or deployment identifier.
  • Validation result and review decision.

A stored run response includes runId at the top level when the run is stored:

{
"status": 200,
"requestId": "req_x1y2z3w4",
"runId": "7KPS70215X0FCDKVQE6HZK4JNA",
"output": {
"data": {
"invoice_number": "INV-2024-0042",
"total_amount": 1547.5
},
"metadata": {},
"pages": [{ "page": 1, "width": 1200, "height": 1697 }]
},
"metrics": { "processingTimeMs": 4800, "pagesProcessed": 1 },
"usage": {
"data_extraction_credits": { "cost": 15, "remainingCredits": 835 }
}
}

Stored runs remain available while the run-storage system retains them. Use requestId for support, and use runId for stored-run retrieval and audit workflows.

Production checklist

Use this checklist before you send extracted data to downstream systems:

  1. Pin x-nutrient-api-version.
  2. Use x-nutrient-engine-version: stable.
  3. Version the schema and extraction instructions.
  4. Hash the schema, instructions, and input document.
  5. Set options.strict: true.
  6. Set options.includeCitations: true.
  7. Set options.multimodal: true only for visual-layout-dependent extraction.
  8. Validate output.data in your application.
  9. Apply per-field confidence thresholds and match-label rules.
  10. Route low-confidence, ungrounded, or missing required fields to manual review.
  11. Retry only 408, 429, 500, and 503 with exponential backoff.
  12. Log requestId, and log runId when you use storeRun.
  13. Use stored processors with pinned published versions instead of latest when Studio-managed configurations need stable behavior.

Next steps

Refer to these guides to continue configuring production extraction: