---
title: "Production structured extraction"
canonical_url: "https://www.nutrient.io/guides/dws-data-extraction/extract/production-structured-extraction/"
md_url: "https://www.nutrient.io/guides/dws-data-extraction/extract/production-structured-extraction.md"
last_updated: "2026-10-02T00:00:00.000Z"
description: "Configure schema-based extraction for production with strict output, citations, confidence gates, validation, retries, and reproducible audit logs."
---

# Production structured extraction

Use the Nutrient DWS Data Extraction API [extract endpoint](https://www.nutrient.io/guides/dws-data-extraction/extract.md) to return structured JSON from documents in production workflows. This guide covers schema validation, citations, confidence gates, retry handling, and reproducible logs.

Use `/extraction/extract` when you know the target fields before processing. For whole-document Markdown or spatial elements, use the [parse endpoint](https://www.nutrient.io/guides/dws-data-extraction/parsing.md) instead.

## Production request template

Start with an explicit API version, the stable engine, a strict schema, citations, and a stored run:

```shell

curl -X POST https://api.nutrient.io/extraction/extract \
  -H "Authorization: Bearer your_api_key_goes_here" \
  -H "x-nutrient-api-version: 2026-05-25" \
  -H "x-nutrient-engine-version: stable" \
  -H "Content-Type: application/json" \
  -d '{
    "url": "https://storage.example.com/invoice.pdf",
    "schema": {
      "type": "object",
      "properties": {
        "invoice_number": {
          "type": "string",
          "description": "Invoice identifier exactly as printed on the invoice."
        },
        "invoice_date": {
          "type": "string",
          "format": "date",
          "description": "Invoice issue date in ISO date format."
        },
        "total_amount": {
          "type": "number",
          "description": "Final invoice total including tax."
        },
        "currency": {
          "type": "string",
          "enum": ["USD", "EUR", "GBP"],
          "description": "Invoice currency as an ISO 4217 code."
        }
      },
      "required": ["invoice_number", "total_amount", "currency"]
    },
    "instructions": "Extract values exactly as shown. Do not infer missing values.",
    "parseConfig": {
      "mode": "understand",
      "options": { "language": "auto" }
    },
    "options": {
      "strict": true,
      "includeCitations": true,
      "multimodal": false
    },
    "storeRun": true
  }'

```

This configuration uses the following settings:

| Setting                                | Purpose                                                                                                                                                 |
| -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `x-nutrient-api-version: 2026-05-25`   | Pins the public API contract for the request.                                                                                                           |
| `x-nutrient-engine-version: stable`    | Uses the stable production engine instead of the `nightly` preview engine.                                                                              |
| `schema`                               | Defines the exact JSON shape to return. The root schema is an object.                                                                                   |
| `options.strict: true`                 | Constrains structured output to the supplied schema.                                                                                                    |
| `options.includeCitations: true`       | Returns per-field grounding metadata in `output.metadata`. Citations are enabled by default, but this setting makes the contract visible.               |
| `parseConfig.mode: "understand"`       | Uses the recommended mode for most forms, invoices, receipts, tables, and key-value extraction workflows.                                               |
| `parseConfig.options.language: "auto"` | Explicitly enables automatic language detection. When you omit `language`, optical character recognition (OCR) also detects the language automatically. |
| `storeRun: true`                       | Stores the run when run storage is available. The response returns a top-level `runId` for audit and debugging workflows.                               |

## Pin versions for reproducibility

Send both version headers on every production request:

```http

x-nutrient-api-version: 2026-05-25
x-nutrient-engine-version: stable

```

Use `stable` for production workflows. Use `nightly` only when you test preview engine behavior before a production rollout.

Log the version headers with each extraction result, because they help you reproduce a run and compare behavior after API or engine changes.

## Define a closed JSON schema

The extract endpoint returns data that matches the JSON Schema you provide. Use an inline schema with a root `object`:

```json

{
  "type": "object",
  "properties": {
    "invoice_number": {
      "type": "string",
      "description": "Invoice identifier exactly as printed on the invoice."
    },
    "invoice_date": {
      "type": "string",
      "format": "date",
      "description": "Invoice issue date in ISO date format."
    },
    "total_amount": {
      "type": "number",
      "description": "Final invoice total including tax."
    },
    "currency": {
      "type": "string",
      "enum": ["USD", "EUR", "GBP"]
    }
  },
  "required": ["invoice_number", "total_amount", "currency"]
}

```

The supported schema keywords include `type`, `properties`, `required`, `items`, `description`, string `enum`, and `format: "date"`. Don’t use unsupported JSON Schema features, such as `$ref`, `$defs`, composition keywords, conditionals, numeric ranges, or string formats other than `date`.

For the full list of supported keywords and limits, refer to the [define a schema](https://www.nutrient.io/guides/dws-data-extraction/extract/define-a-schema.md) guide.

## Use strict output and application validation

Set `options.strict` to `true` for production schema extraction:

```json

{
  "options": {
    "strict": true,
    "includeCitations": true
  }
}

```

`strict: true` constrains the extraction model to emit JSON that conforms to the supplied schema. It improves generation behavior, but it doesn’t replace validation in your application.

Validate `output.data` with the same schema before you use the result downstream, and then apply business rules that aren’t part of the supported schema subset. These rules can include numeric ranges, identifier formats, currency rules, or cross-field checks:

```javascript

import Ajv from "ajv";

const ajv = new Ajv();
const validate = ajv.compile(schema);

if (!validate(result.output.data)) {
  throw new Error(`Schema validation failed: ${ajv.errorsText(validate.errors)}`);
}

```

## Choose the parse mode

The extract endpoint runs a parse stage before schema extraction, and you configure that stage with `parseConfig.mode`.

| Mode         | Use when                                                                                        |
| ------------ | ----------------------------------------------------------------------------------------------- |
| `structure`  | Documents have clean, predictable layouts, and lower cost or lower latency matters most.        |
| `understand` | Most invoices, forms, receipts, tables, and key-value extraction workflows.                     |
| `agentic`    | Documents require complex visual reasoning, such as degraded scans, dense layouts, or diagrams. |

Start with `understand` for most structured extraction. Move to `structure` after you test a representative sample and confirm the output quality. Move to `agentic` when `understand` misses visual context or degraded content.

Set `language` to `"auto"` when you want automatic OCR language detection:

```json

{
  "parseConfig": {
    "mode": "understand",
    "options": { "language": "auto" }
  }
}

```

When you omit `language`, OCR detects the language automatically. For known languages, use a language name, an ISO 639-2 code, an array, or a `+`-joined string. For details, refer to the [parse configuration](https://www.nutrient.io/guides/dws-data-extraction/extract/parse-configuration.md) and [supported languages](https://www.nutrient.io/guides/dws-data-extraction/supported-languages.md) guides.

## Enable multimodal extraction only when needed

Set `options.multimodal` to `true` when the extraction depends on visual layout or page images:

```json

{
  "options": {
    "multimodal": true
  }
}

```

Multimodal extraction sends rendered page images to the extraction model with parsed text. Use it for visually complex layouts, image-heavy documents, difficult forms, or fields whose meaning depends on spatial layout.

Multimodal extraction increases cost and latency. Keep it disabled for text-centric documents when parsed text and layout structure provide enough context.

## Gate results with citations and confidence

With `options.includeCitations: true`, `output.metadata` mirrors `output.data`. Each scalar field can include a citation with a `match` label, page reference, bounding box, and confidence score:

```json

{
  "output": {
    "data": {
      "invoice_number": "INV-2024-0042",
      "total_amount": 1547.5
    },
    "metadata": {
      "invoice_number": {
        "match": "id_match",
        "confidence": 0.93,
        "pageNumber": 1,
        "bbox": { "x": 878, "y": 268, "width": 82, "height": 25 }
      },
      "total_amount": {
        "match": "fuzzy_match",
        "confidence": 0.76,
        "pageNumber": 1,
        "bbox": { "x": 930, "y": 1200, "width": 96, "height": 28 }
      }
    }
  }
}

```

Use both `match` and `confidence` in review logic:

- Route `fuzzy_match` and `not_found` fields to review.

- Treat missing citations for required fields as review or failure.

- Set per-field thresholds. Identifiers, totals, and regulated fields usually need stricter thresholds than descriptive text.

Confidence values are relative signals, not calibrated probabilities. Calibrate thresholds on your own labeled sample set for each document type and field.

Use the following policy as a starting point for calibration:

| Confidence       | Example action                          |
| ---------------- | --------------------------------------- |
| `>= 0.90`        | Accept automatically.                   |
| `0.70` to `0.89` | Accept with warning or route to review. |
| `< 0.70`         | Reject or route to manual review.       |

Implement the policy in application code:

```javascript

const REQUIRED_CONFIDENCE = {
  invoice_number: 0.9,
  total_amount: 0.9,
  currency: 0.85,
};

function assertConfidence(result) {
  const metadata = result.output.metadata?? {};

  for (const [field, minConfidence] of Object.entries(REQUIRED_CONFIDENCE)) {
    const citation = metadata[field];

    if (!citation) {
      throw new Error(`Missing citation for ${field}`);
    }

    if (citation.match === "fuzzy_match" || citation.match === "not_found") {
      throw new Error(`Review required for ${field}: ${citation.match}`);
    }

    if (citation.confidence!= null && citation.confidence < minConfidence) {
      throw new Error(`Low confidence for ${field}`);
    }
  }
}

```

For nested objects and arrays, traverse `output.data` and `output.metadata` together. For traversal examples, refer to the [citations and confidence](https://www.nutrient.io/guides/dws-data-extraction/extract/citations-and-confidence.md) guide.

## Handle errors explicitly

Branch on the HTTP status code, and retry only transient errors.

| Status | Meaning                                   | Handling                        |
| ------ | ----------------------------------------- | ------------------------------- |
| 400    | Bad request, invalid schema, invalid mode | Don’t retry. Fix the request.   |
| 401    | Invalid or missing API key                | Don’t retry. Fix credentials.   |
| 402    | Insufficient credits                      | Don’t retry automatically.      |
| 408    | Timeout                                   | Retry with exponential backoff. |
| 413    | File too large                            | Split, compress, or reject.     |
| 422    | Remote URL rejected or unavailable        | Fix the URL or input source.    |
| 429    | Rate limited                              | Retry with exponential backoff. |
| 500    | Internal processing error                 | Retry and keep the `requestId`. |
| 503    | Backend unavailable or overloaded         | Retry with exponential backoff. |

Use bounded retries with jitter for `408`, `429`, `500`, and `503`. Keep `requestId` in logs for support and debugging.

For error response shapes and troubleshooting details, refer to the [error handling](https://www.nutrient.io/guides/dws-data-extraction/errors.md) guide.

## Log audit and reproducibility fields

Store enough information to reproduce and debug each extraction, but avoid logging sensitive document content unless you need it.

Log these fields for every production extraction:

- `requestId`.

- `runId`, when `storeRun: true` returns it.

- `x-nutrient-api-version`.

- `x-nutrient-engine-version`.

- Parse mode and parse options.

- Extract options, including `strict`, `includeCitations`, and `multimodal`.

- Schema version and schema hash.

- Instruction version and instruction hash.

- Input document hash.

- Application version or deployment identifier.

- Validation result and review decision.

A stored run response includes `runId` at the top level when the run is stored:

```json

{
  "status": 200,
  "requestId": "req_x1y2z3w4",
  "runId": "7KPS70215X0FCDKVQE6HZK4JNA",
  "output": {
    "data": {
      "invoice_number": "INV-2024-0042",
      "total_amount": 1547.5
    },
    "metadata": {},
    "pages": [{ "page": 1, "width": 1200, "height": 1697 }]
  },
  "metrics": { "processingTimeMs": 4800, "pagesProcessed": 1 },
  "usage": {
    "data_extraction_credits": { "cost": 15, "remainingCredits": 835 }
  }
}

```

Stored runs remain available while the run-storage system retains them. Use `requestId` for support, and use `runId` for stored-run retrieval and audit workflows.

## Production checklist

Use this checklist before you send extracted data to downstream systems:

1. Pin `x-nutrient-api-version`.

2. Use `x-nutrient-engine-version: stable`.

3. Version the schema and extraction instructions.

4. Hash the schema, instructions, and input document.

5. Set `options.strict: true`.

6. Set `options.includeCitations: true`.

7. Set `options.multimodal: true` only for visual-layout-dependent extraction.

8. Validate `output.data` in your application.

9. Apply per-field confidence thresholds and match-label rules.

10. Route low-confidence, ungrounded, or missing required fields to manual review.

11. Retry only `408`, `429`, `500`, and `503` with exponential backoff.

12. Log `requestId`, and log `runId` when you use `storeRun`.

13. Use stored processors with pinned published versions instead of `latest` when Studio-managed configurations need stable behavior.

## Next steps

Refer to these guides to continue configuring production extraction:

- [Define a schema](https://www.nutrient.io/guides/dws-data-extraction/extract/define-a-schema.md) for schema keywords and limits.

- [Parse configuration](https://www.nutrient.io/guides/dws-data-extraction/extract/parse-configuration.md) for parse modes and language settings.

- [Citations and confidence](https://www.nutrient.io/guides/dws-data-extraction/extract/citations-and-confidence.md) for grounding metadata and review workflows.

- [Error handling](https://www.nutrient.io/guides/dws-data-extraction/errors.md) for response formats and troubleshooting.
---

## Related pages

- [Extract endpoint](/guides/dws-data-extraction/extract.md)
- [Citations and confidence](/guides/dws-data-extraction/extract/citations-and-confidence.md)
- [Define a schema](/guides/dws-data-extraction/extract/define-a-schema.md)
- [Parse configuration](/guides/dws-data-extraction/extract/parse-configuration.md)

