Production structured extraction
Use the Nutrient DWS Data Extraction API extract endpoint to return structured JSON from documents in production workflows. This guide covers schema validation, citations, confidence gates, retry handling, and reproducible logs.
Use /extraction/extract when you know the target fields before processing. For whole-document Markdown or spatial elements, use the parse endpoint instead.
Production request template
Start with an explicit API version, the stable engine, a strict schema, citations, and a stored run:
curl -X POST https://api.nutrient.io/extraction/extract \ -H "Authorization: Bearer your_api_key_goes_here" \ -H "x-nutrient-api-version: 2026-05-25" \ -H "x-nutrient-engine-version: stable" \ -H "Content-Type: application/json" \ -d '{ "url": "https://storage.example.com/invoice.pdf", "schema": { "type": "object", "properties": { "invoice_number": { "type": "string", "description": "Invoice identifier exactly as printed on the invoice." }, "invoice_date": { "type": "string", "format": "date", "description": "Invoice issue date in ISO date format." }, "total_amount": { "type": "number", "description": "Final invoice total including tax." }, "currency": { "type": "string", "enum": ["USD", "EUR", "GBP"], "description": "Invoice currency as an ISO 4217 code." } }, "required": ["invoice_number", "total_amount", "currency"] }, "instructions": "Extract values exactly as shown. Do not infer missing values.", "parseConfig": { "mode": "understand", "options": { "language": "auto" } }, "options": { "strict": true, "includeCitations": true, "multimodal": false }, "storeRun": true }'This configuration uses the following settings:
| Setting | Purpose |
|---|---|
x-nutrient-api-version: 2026-05-25 | Pins the public API contract for the request. |
x-nutrient-engine-version: stable | Uses the stable production engine instead of the nightly preview engine. |
schema | Defines the exact JSON shape to return. The root schema is an object. |
options.strict: true | Constrains structured output to the supplied schema. |
options.includeCitations: true | Returns per-field grounding metadata in output.metadata. Citations are enabled by default, but this setting makes the contract visible. |
parseConfig.mode: "understand" | Uses the recommended mode for most forms, invoices, receipts, tables, and key-value extraction workflows. |
parseConfig.options.language: "auto" | Explicitly enables automatic language detection. When you omit language, optical character recognition (OCR) also detects the language automatically. |
storeRun: true | Stores the run when run storage is available. The response returns a top-level runId for audit and debugging workflows. |
Pin versions for reproducibility
Send both version headers on every production request:
x-nutrient-api-version: 2026-05-25x-nutrient-engine-version: stableUse stable for production workflows. Use nightly only when you test preview engine behavior before a production rollout.
Log the version headers with each extraction result, because they help you reproduce a run and compare behavior after API or engine changes.
Define a closed JSON schema
The extract endpoint returns data that matches the JSON Schema you provide. Use an inline schema with a root object:
{ "type": "object", "properties": { "invoice_number": { "type": "string", "description": "Invoice identifier exactly as printed on the invoice." }, "invoice_date": { "type": "string", "format": "date", "description": "Invoice issue date in ISO date format." }, "total_amount": { "type": "number", "description": "Final invoice total including tax." }, "currency": { "type": "string", "enum": ["USD", "EUR", "GBP"] } }, "required": ["invoice_number", "total_amount", "currency"]}The supported schema keywords include type, properties, required, items, description, string enum, and format: "date". Don’t use unsupported JSON Schema features, such as $ref, $defs, composition keywords, conditionals, numeric ranges, or string formats other than date.
For the full list of supported keywords and limits, refer to the define a schema guide.
Use strict output and application validation
Set options.strict to true for production schema extraction:
{ "options": { "strict": true, "includeCitations": true }}strict: true constrains the extraction model to emit JSON that conforms to the supplied schema. It improves generation behavior, but it doesn’t replace validation in your application.
Validate output.data with the same schema before you use the result downstream, and then apply business rules that aren’t part of the supported schema subset. These rules can include numeric ranges, identifier formats, currency rules, or cross-field checks:
import Ajv from "ajv";
const ajv = new Ajv();const validate = ajv.compile(schema);
if (!validate(result.output.data)) { throw new Error(`Schema validation failed: ${ajv.errorsText(validate.errors)}`);}Choose the parse mode
The extract endpoint runs a parse stage before schema extraction, and you configure that stage with parseConfig.mode.
| Mode | Use when |
|---|---|
structure | Documents have clean, predictable layouts, and lower cost or lower latency matters most. |
understand | Most invoices, forms, receipts, tables, and key-value extraction workflows. |
agentic | Documents require complex visual reasoning, such as degraded scans, dense layouts, or diagrams. |
Start with understand for most structured extraction. Move to structure after you test a representative sample and confirm the output quality. Move to agentic when understand misses visual context or degraded content.
Set language to "auto" when you want automatic OCR language detection:
{ "parseConfig": { "mode": "understand", "options": { "language": "auto" } }}When you omit language, OCR detects the language automatically. For known languages, use a language name, an ISO 639-2 code, an array, or a +-joined string. For details, refer to the parse configuration and supported languages guides.
Enable multimodal extraction only when needed
Set options.multimodal to true when the extraction depends on visual layout or page images:
{ "options": { "multimodal": true }}Multimodal extraction sends rendered page images to the extraction model with parsed text. Use it for visually complex layouts, image-heavy documents, difficult forms, or fields whose meaning depends on spatial layout.
Multimodal extraction increases cost and latency. Keep it disabled for text-centric documents when parsed text and layout structure provide enough context.
Gate results with citations and confidence
With options.includeCitations: true, output.metadata mirrors output.data. Each scalar field can include a citation with a match label, page reference, bounding box, and confidence score:
{ "output": { "data": { "invoice_number": "INV-2024-0042", "total_amount": 1547.5 }, "metadata": { "invoice_number": { "match": "id_match", "confidence": 0.93, "pageNumber": 1, "bbox": { "x": 878, "y": 268, "width": 82, "height": 25 } }, "total_amount": { "match": "fuzzy_match", "confidence": 0.76, "pageNumber": 1, "bbox": { "x": 930, "y": 1200, "width": 96, "height": 28 } } } }}Use both match and confidence in review logic:
- Route
fuzzy_matchandnot_foundfields to review. - Treat missing citations for required fields as review or failure.
- Set per-field thresholds. Identifiers, totals, and regulated fields usually need stricter thresholds than descriptive text.
Confidence values are relative signals, not calibrated probabilities. Calibrate thresholds on your own labeled sample set for each document type and field.
Use the following policy as a starting point for calibration:
| Confidence | Example action |
|---|---|
>= 0.90 | Accept automatically. |
0.70 to 0.89 | Accept with warning or route to review. |
< 0.70 | Reject or route to manual review. |
Implement the policy in application code:
const REQUIRED_CONFIDENCE = { invoice_number: 0.9, total_amount: 0.9, currency: 0.85,};
function assertConfidence(result) { const metadata = result.output.metadata ?? {};
for (const [field, minConfidence] of Object.entries(REQUIRED_CONFIDENCE)) { const citation = metadata[field];
if (!citation) { throw new Error(`Missing citation for ${field}`); }
if (citation.match === "fuzzy_match" || citation.match === "not_found") { throw new Error(`Review required for ${field}: ${citation.match}`); }
if (citation.confidence != null && citation.confidence < minConfidence) { throw new Error(`Low confidence for ${field}`); } }}For nested objects and arrays, traverse output.data and output.metadata together. For traversal examples, refer to the citations and confidence guide.
Handle errors explicitly
Branch on the HTTP status code, and retry only transient errors.
| Status | Meaning | Handling |
|---|---|---|
| 400 | Bad request, invalid schema, invalid mode | Don’t retry. Fix the request. |
| 401 | Invalid or missing API key | Don’t retry. Fix credentials. |
| 402 | Insufficient credits | Don’t retry automatically. |
| 408 | Timeout | Retry with exponential backoff. |
| 413 | File too large | Split, compress, or reject. |
| 422 | Remote URL rejected or unavailable | Fix the URL or input source. |
| 429 | Rate limited | Retry with exponential backoff. |
| 500 | Internal processing error | Retry and keep the requestId. |
| 503 | Backend unavailable or overloaded | Retry with exponential backoff. |
Use bounded retries with jitter for 408, 429, 500, and 503. Keep requestId in logs for support and debugging.
For error response shapes and troubleshooting details, refer to the error handling guide.
Log audit and reproducibility fields
Store enough information to reproduce and debug each extraction, but avoid logging sensitive document content unless you need it.
Log these fields for every production extraction:
requestId.runId, whenstoreRun: truereturns it.x-nutrient-api-version.x-nutrient-engine-version.- Parse mode and parse options.
- Extract options, including
strict,includeCitations, andmultimodal. - Schema version and schema hash.
- Instruction version and instruction hash.
- Input document hash.
- Application version or deployment identifier.
- Validation result and review decision.
A stored run response includes runId at the top level when the run is stored:
{ "status": 200, "requestId": "req_x1y2z3w4", "runId": "7KPS70215X0FCDKVQE6HZK4JNA", "output": { "data": { "invoice_number": "INV-2024-0042", "total_amount": 1547.5 }, "metadata": {}, "pages": [{ "page": 1, "width": 1200, "height": 1697 }] }, "metrics": { "processingTimeMs": 4800, "pagesProcessed": 1 }, "usage": { "data_extraction_credits": { "cost": 15, "remainingCredits": 835 } }}Stored runs remain available while the run-storage system retains them. Use requestId for support, and use runId for stored-run retrieval and audit workflows.
Production checklist
Use this checklist before you send extracted data to downstream systems:
- Pin
x-nutrient-api-version. - Use
x-nutrient-engine-version: stable. - Version the schema and extraction instructions.
- Hash the schema, instructions, and input document.
- Set
options.strict: true. - Set
options.includeCitations: true. - Set
options.multimodal: trueonly for visual-layout-dependent extraction. - Validate
output.datain your application. - Apply per-field confidence thresholds and match-label rules.
- Route low-confidence, ungrounded, or missing required fields to manual review.
- Retry only
408,429,500, and503with exponential backoff. - Log
requestId, and logrunIdwhen you usestoreRun. - Use stored processors with pinned published versions instead of
latestwhen Studio-managed configurations need stable behavior.
Next steps
Refer to these guides to continue configuring production extraction:
- Define a schema for schema keywords and limits.
- Parse configuration for parse modes and language settings.
- Citations and confidence for grounding metadata and review workflows.
- Error handling for response formats and troubleshooting.