---
title: "Extract structured data in C# | Nutrient .NET SDK"
canonical_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/extract-structured-data/"
md_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/extract-structured-data.md"
last_updated: "2026-10-08T00:00:00.000Z"
description: "Extract schema-shaped JSON data from documents using Nutrient .NET SDK."
---

# Extracting structured data from documents

Most document workflows don't want a wall of recognized text — they want *fields*: the invoice total, the patient's date of birth, every line item as a row. *Structured extraction* turns a document into exactly the JSON you ask for: you supply a JSON Schema describing the fields, and an AI model fills it from the document's recognized content.

This sample shows how to extract schema-shaped data from a document using Nutrient.NET SDK. The result reports not just the values but also *where* each value came from — per-field source locations and grounding labels you can use to verify the extraction against the original document.

[Download sample](https://www.nutrient.io/downloads/samples/csharp/extract-structured-data.zip)

## How Nutrient helps

The `DataExtraction` class runs the full structured extraction workflow behind a single method call. It runs the same extract operation, under the same option names, and returns the same output files as the hosted Nutrient data extraction service, so the code and the option names you learn here carry over between the two. The SDK handles:

- Reading the document and recognizing text, tables, key-value regions, and form fields in reading order

- Sending the recognized content and your JSON Schema to the AI model as a structured-output request

- Retrying automatically when the model's response doesn't conform to the schema

- Grounding each extracted value back to its source location in the document

- Serializing the result to JSON

A successful extraction conforms to your schema — the same call with the same schema yields the same shape, ready for your downstream code to consume without defensive parsing. If the model exhausts its retries, the default behavior is to return an empty extraction with a warning. The sample instead turns that fallback off so a failed extraction throws rather than looking like valid business data.

## Parse isn't a prerequisite

`Extract` and `Parse` are two independent operations on the same class. `Extract` reads the document itself: it runs the extraction pipeline internally, feeds the recognized content to the model, and returns the filled schema. You don't have to call `Parse` first, and passing parsed output into `Extract` isn't a supported input. Call `Parse` when you want the document's structure as JSON or Markdown; call `Extract` when you want your schema's fields filled in.

## What shapes the result

Two inputs shape the result:

- **JSON Schema** (required) — The fields to extract. Supply it inline with `ExtractionSchema` or from a file with `ExtractionSchemaFile`; the two are mutually exclusive. Each schema property's `description` tells the extractor what belongs there — the better the description, the better the match.

- **Instructions** (optional) — Free-form guidance for the extraction: disambiguation rules, formatting preferences, or domain context — anything you'd tell a colleague doing the extraction by hand.

Extraction requires an AI model. Configure the provider on the document's `ExtractSettings`; it's the same value as `AiProcessingSettings`, so a provider set for the whole SDK applies too. A local OpenAI-compatible server keeps documents on your machine; a hosted provider needs an API key. Structured extraction requires the vision data extraction feature in your license.

## Complete implementation

Import the Nutrient namespace:

```csharp

using Nutrient;

```

## Opening the document

Open the source document and bind a `DataExtraction` instance to it. Both are disposable, so use C#’s [using statement](https://learn.microsoft.com/dotnet/csharp/language-reference/statements/using) for each — the document releases the file, and the `DataExtraction` instance drops its reference so the native bridge can free the handle that pins the document:

```csharp

try
{
    using Document document = Document.Open("input_invoice_lumen.pdf");
    using DataExtraction dataExtraction = DataExtraction.Set(document);

```

## Supplying the schema

Like every SDK operation, `Extract()` reads its options from the document settings. `ExtractSettings` holds the options of the hosted extraction service under the same names; options you don't set keep the default the extraction pipeline or the SDK-wide settings provide.

This sample ships `input_invoice_lumen_schema.json`, a JSON Schema describing the invoice fields to pull out. Point `ExtractionSchemaFile` at it rather than embedding the schema in your source, so the schema stays editable without a rebuild.

The settings wrapper holds a native handle, so declare it with `using` as well. Ask for a failed extraction to throw instead of returning an empty result:

```csharp

    using ExtractSettings settings = document.Settings.ExtractSettings;
    settings.ExtractionSchemaFile = "input_invoice_lumen_schema.json";
    settings.Instructions = "Amounts are plain numbers without currency symbols.";
    document.Settings.AiProcessingSettings.UnparseableOutputHandling = UnparseableOutputHandling.Fail;

```

The bundled schema is a plain JSON Schema object — `type`, `properties`, and `required` — where every property carries the `description` the extractor matches against the document:

```json

{
  "type": "object",
  "properties": {
    "invoice number": {
      "type": "string",
      "description": "The unique invoice number or identifier"
    },
    "total": {
      "type": "number",
      "description": "The final total amount of the invoice including all taxes and discounts."
    }
  },
  "required": ["invoice number", "total"]
}

```

To build the schema at runtime instead, assign the same JSON to `settings.ExtractionSchema` as a string.

## Configuring the AI provider

Point the run at your model. This example uses a local OpenAI-compatible endpoint, so the document never leaves your machine:

```csharp

    settings.Provider = "local";
    settings.Endpoint = "http://localhost:1234/v1";
    settings.Model = "your-model-id";

```

For a hosted provider instead, set `Provider` to `"openai"` or `"anthropic"` along with `ApiKey` (and `Endpoint` for an OpenAI-compatible proxy).

## Strict structured output

Strict structured output is enabled by default through the document settings. The sample assigns it explicitly so the behavior is clear: the model's response is grammar-constrained to the schema, and a successful result matches it exactly.

```csharp

    settings.StrictStructuredOutput = true;

```

Two things to know about strict mode:

- **You don't change your schema.** Strict mode has formal requirements (every object closed, every property accounted for), and the SDK normalizes your schema to satisfy them automatically.

- **Absent fields come back as `null`.** In strict mode the model must emit every schema property, so a field the document doesn't contain is returned as `null` instead of being omitted — your downstream code can rely on every key being present. Fields you list in the schema's `required` array keep their declared type untouched — they can only be `null` if your own schema allows it.

Strict mode requires a model and endpoint that support grammar-constrained structured outputs. Check your provider or local server's documentation. If it doesn't support them, set `StrictStructuredOutput` to `false`; the SDK then validates the model response and retries when it doesn't conform.

## Extracting the data

Call `Extract()` and read the `result.json` artifact. Every run produces it, so it's safe to read by name. The result wrapper holds native resources as well, so it gets a `using` declaration too — read what you need from it before it goes out of scope at the end of the block:

```csharp

    using ExtractResult result = dataExtraction.Extract();
    File.WriteAllText("output.json", result.GetArtifact("result.json"));
    Console.WriteLine("Extraction written to output.json");

```

## Handling the conditional artifacts

Besides `result.json`, a run can emit up to three more artifacts, each only when the conditions for it are met. Check with `HasArtifact` before reading, because `GetArtifact` throws on a name the run didn't produce:

```csharp

    if (result.HasArtifact("warnings.json"))
    {
        Console.WriteLine($"Warnings: {result.GetArtifact("warnings.json")}");
    }

    Console.WriteLine($"Artifacts: {string.Join(", ", result.ArtifactNames)}");
}

```

| Artifact           | Emitted when                                                          |
| ------------------ | --------------------------------------------------------------------- |
| `result.json`      | Always.                                                               |
| `warnings.json`    | At least one non-fatal AI-layer warning was collected.                |
| `confidence.json`  | `IncludeConfidence` is true and at least one scoreable signal exists. |
| `diagnostics.json` | `IncludeDiagnostics` is true.                                         |

`result.SaveTo("artifacts")` writes every artifact the run produced into a directory, each under its own name — useful when you want the whole set on disk without checking each one.

## Understanding the output

`result.json` has two top-level nodes:

- **`extraction`** — The extracted fields, shaped exactly to your schema.

- **`metadata`** — One entry per extracted field, carrying where the value came from: a `match` grounding label and source location info (page and bounding box) so you can highlight the source in a viewer or route low-trust fields to human review.

The `match` label tells you how confidently the value was traced back to the document: an exact source match, a partial or multi-block match, a fuzzy match, or `not_found` when the value couldn't be located in the recognized content — the strongest signal that a field deserves review.

Source grounding is on by default. When only the extracted values matter, turn it off with `settings.IncludeSourceLocations = false` to cut model token usage — grounding asks the model to also return per-field source references, which roughly doubles the schema sent with each request.

For per-field confidence signals, set `settings.IncludeConfidence = true`. Each metadata entry then also carries the individual confidence components for the field, and the run emits the `confidence.json` sidecar. Combined with the `match` grounding labels, this gives your pipeline a per-field basis for automatic acceptance versus human review.

## Error handling

The SDK throws a `NutrientException` when extraction fails:

```csharp

catch (NutrientException e)
{
    Console.Error.WriteLine($"Error: {e.Message}");
    Environment.Exit(1);
}

```

Common failure scenarios include:

- The document can't be read due to path or permission issues

- The JSON Schema is missing or malformed (validated before any model call)

- The model endpoint is unreachable, the model can't produce a schema-conforming response, or the feature isn't licensed

In production code:

- Catch `NutrientException`.

- Return a clear error message.

- Log failure details for debugging.

## Conclusion

The workflow for structured data extraction is:

1. Open the source document, and bind a `DataExtraction` instance to it with `DataExtraction.Set()` — both in a [using statement](https://learn.microsoft.com/dotnet/csharp/language-reference/statements/using).

2. Point `ExtractionSchemaFile` (or `ExtractionSchema`) at the JSON Schema to fill, and add optional instructions.

3. Configure the AI provider on the document settings — local for privacy, or a hosted provider — or inherit it from the SDK-wide settings.

4. Call `Extract()` and read the `result.json` artifact; consume the `extraction` node and use `metadata` to verify or review.

5. Check `HasArtifact` for the conditional artifacts, or call `SaveTo()` to write the whole set.

6. Handle `NutrientException` for robust error recovery.

For related extraction workflows, refer to the [.NET SDK guides](https://www.nutrient.io/guides/dotnet/csharp.md).

Download [this ready-to-use sample package](https://www.nutrient.io/downloads/samples/csharp/extract-structured-data.zip) to explore structured data extraction.
---

## Related pages

- [Nutrient .NET SDK extraction guides](/guides/dotnet/csharp/extraction.md)
- [Applying OCR to a PDF page](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf-page.md)
- [Applying OCR to a PDF document](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf.md)
- [Classifying documents](/guides/dotnet/csharp/extraction/classify-document.md)
- [Generating image descriptions using Claude](/guides/dotnet/csharp/extraction/describe-image-with-claude.md)
- [Generating image descriptions using local AI](/guides/dotnet/csharp/extraction/describe-image-with-local-ai.md)
- [Generating image descriptions using OpenAI](/guides/dotnet/csharp/extraction/describe-image-with-openai.md)
- [Detecting document language](/guides/dotnet/csharp/extraction/detect-document-language.md)
- [Extracting data from images using ICR](/guides/dotnet/csharp/extraction/extract-data-from-image-icr.md)
- [Extracting data from images using OCR](/guides/dotnet/csharp/extraction/extract-data-from-image-ocr.md)
- [Extracting data from images using vision language models](/guides/dotnet/csharp/extraction/extract-data-from-image-vlm.md)
- [Extracting data from specific pages](/guides/dotnet/csharp/extraction/extract-data-from-specific-pages.md)
- [Extracting form fields from images](/guides/dotnet/csharp/extraction/extract-form-fields-from-image.md)
- [Generating extraction schemas](/guides/dotnet/csharp/extraction/generate-extraction-schema.md)
- [Extracting JSON data from a PDF document](/guides/dotnet/csharp/extraction/json-data-extraction.md)
- [Labeling form fields with a vision language model](/guides/dotnet/csharp/extraction/label-form-fields-with-vlm.md)
- [Parsing a document into structured content](/guides/dotnet/csharp/extraction/parse-document.md)
- [Extracting text from PDF documents](/guides/dotnet/csharp/extraction/pdf-to-text.md)
- [Reading barcodes with vision extraction](/guides/dotnet/csharp/extraction/read-barcodes-with-vision.md)
- [Extracting text from multilingual images](/guides/dotnet/csharp/extraction/read-text-from-image-multi-language.md)
- [Extracting text from images](/guides/dotnet/csharp/extraction/read-text-from-image.md)
- [Searching document text](/guides/dotnet/csharp/extraction/search-document-text.md)
- [Speeding up first ICR operation by predownloading models](/guides/dotnet/csharp/extraction/speed-up-first-icr-by-downloading-requirements.md)

