This HTML page is not optimized for LLM or AI agent consumption. Fetch the Markdown version instead: /guides/dotnet/csharp/extraction/extract-structured-data.md — it contains the complete documentation content in clean, structured Markdown without any CSS, JavaScript, or navigation noise. Extract structured data in C# | Nutrient .NET SDK

Most document workflows don’t want a wall of recognized text — they want fields: the invoice total, the patient’s date of birth, every line item as a row. Structured extraction turns a document into exactly the JSON you ask for: you supply a JSON Schema describing the fields, and an AI model fills it from the document’s recognized content.

This sample shows how to extract schema-shaped data from a document using Nutrient .NET SDK. The result reports not just the values but also where each value came from — per-field source locations and grounding labels you can use to verify the extraction against the original document.

Download sample

How Nutrient helps

The DataExtraction class runs the full structured extraction workflow behind a single method call. It runs the same extract operation, under the same option names, and returns the same output files as the hosted Nutrient data extraction service, so the code and the option names you learn here carry over between the two. The SDK handles:

  • Reading the document and recognizing text, tables, key-value regions, and form fields in reading order
  • Sending the recognized content and your JSON Schema to the AI model as a structured-output request
  • Retrying automatically when the model’s response doesn’t conform to the schema
  • Grounding each extracted value back to its source location in the document
  • Serializing the result to JSON

A successful extraction conforms to your schema — the same call with the same schema yields the same shape, ready for your downstream code to consume without defensive parsing. If the model exhausts its retries, the default behavior is to return an empty extraction with a warning. The sample instead turns that fallback off so a failed extraction throws rather than looking like valid business data.

Parse isn’t a prerequisite

Extract and Parse are two independent operations on the same class. Extract reads the document itself: it runs the extraction pipeline internally, feeds the recognized content to the model, and returns the filled schema. You don’t have to call Parse first, and passing parsed output into Extract isn’t a supported input. Call Parse when you want the document’s structure as JSON or Markdown; call Extract when you want your schema’s fields filled in.

What shapes the result

Two inputs shape the result:

  • JSON Schema (required) — The fields to extract. Supply it inline with ExtractionSchema or from a file with ExtractionSchemaFile; the two are mutually exclusive. Each schema property’s description tells the extractor what belongs there — the better the description, the better the match.
  • Instructions (optional) — Free-form guidance for the extraction: disambiguation rules, formatting preferences, or domain context — anything you’d tell a colleague doing the extraction by hand.

Extraction requires an AI model. Configure the provider on the document’s ExtractSettings; it’s the same value as AiProcessingSettings, so a provider set for the whole SDK applies too. A local OpenAI-compatible server keeps documents on your machine; a hosted provider needs an API key. Structured extraction requires the vision data extraction feature in your license.

Complete implementation

Import the Nutrient namespace:

using Nutrient;

Opening the document

Open the source document and bind a DataExtraction instance to it. Both are disposable, so use C#’s using statement(opens in a new tab) for each — the document releases the file, and the DataExtraction instance drops its reference so the native bridge can free the handle that pins the document:

try
{
using Document document = Document.Open("input_invoice_lumen.pdf");
using DataExtraction dataExtraction = DataExtraction.Set(document);

Supplying the schema

Like every SDK operation, Extract() reads its options from the document settings. ExtractSettings holds the options of the hosted extraction service under the same names; options you don’t set keep the default the extraction pipeline or the SDK-wide settings provide.

This sample ships input_invoice_lumen_schema.json, a JSON Schema describing the invoice fields to pull out. Point ExtractionSchemaFile at it rather than embedding the schema in your source, so the schema stays editable without a rebuild.

The settings wrapper holds a native handle, so declare it with using as well. Ask for a failed extraction to throw instead of returning an empty result:

using ExtractSettings settings = document.Settings.ExtractSettings;
settings.ExtractionSchemaFile = "input_invoice_lumen_schema.json";
settings.Instructions = "Amounts are plain numbers without currency symbols.";
document.Settings.AiProcessingSettings.UnparseableOutputHandling = UnparseableOutputHandling.Fail;

The bundled schema is a plain JSON Schema object — type, properties, and required — where every property carries the description the extractor matches against the document:

{
"type": "object",
"properties": {
"invoice number": {
"type": "string",
"description": "The unique invoice number or identifier"
},
"total": {
"type": "number",
"description": "The final total amount of the invoice including all taxes and discounts."
}
},
"required": ["invoice number", "total"]
}

To build the schema at runtime instead, assign the same JSON to settings.ExtractionSchema as a string.

Configuring the AI provider

Point the run at your model. This example uses a local OpenAI-compatible endpoint, so the document never leaves your machine:

settings.Provider = "local";
settings.Endpoint = "http://localhost:1234/v1";
settings.Model = "your-model-id";

For a hosted provider instead, set Provider to "openai" or "anthropic" along with ApiKey (and Endpoint for an OpenAI-compatible proxy).

Strict structured output

Strict structured output is enabled by default through the document settings. The sample assigns it explicitly so the behavior is clear: the model’s response is grammar-constrained to the schema, and a successful result matches it exactly.

settings.StrictStructuredOutput = true;

Two things to know about strict mode:

  • You don’t change your schema. Strict mode has formal requirements (every object closed, every property accounted for), and the SDK normalizes your schema to satisfy them automatically.
  • Absent fields come back as null. In strict mode the model must emit every schema property, so a field the document doesn’t contain is returned as null instead of being omitted — your downstream code can rely on every key being present. Fields you list in the schema’s required array keep their declared type untouched — they can only be null if your own schema allows it.

Strict mode requires a model and endpoint that support grammar-constrained structured outputs. Check your provider or local server’s documentation. If it doesn’t support them, set StrictStructuredOutput to false; the SDK then validates the model response and retries when it doesn’t conform.

Extracting the data

Call Extract() and read the result.json artifact. Every run produces it, so it’s safe to read by name. The result wrapper holds native resources as well, so it gets a using declaration too — read what you need from it before it goes out of scope at the end of the block:

using ExtractResult result = dataExtraction.Extract();
File.WriteAllText("output.json", result.GetArtifact("result.json"));
Console.WriteLine("Extraction written to output.json");

Handling the conditional artifacts

Besides result.json, a run can emit up to three more artifacts, each only when the conditions for it are met. Check with HasArtifact before reading, because GetArtifact throws on a name the run didn’t produce:

if (result.HasArtifact("warnings.json"))
{
Console.WriteLine($"Warnings: {result.GetArtifact("warnings.json")}");
}
Console.WriteLine($"Artifacts: {string.Join(", ", result.ArtifactNames)}");
}
ArtifactEmitted when
result.jsonAlways.
warnings.jsonAt least one non-fatal AI-layer warning was collected.
confidence.jsonIncludeConfidence is true and at least one scoreable signal exists.
diagnostics.jsonIncludeDiagnostics is true.

result.SaveTo("artifacts") writes every artifact the run produced into a directory, each under its own name — useful when you want the whole set on disk without checking each one.

Understanding the output

result.json has two top-level nodes:

  • extraction — The extracted fields, shaped exactly to your schema.
  • metadata — One entry per extracted field, carrying where the value came from: a match grounding label and source location info (page and bounding box) so you can highlight the source in a viewer or route low-trust fields to human review.

The match label tells you how confidently the value was traced back to the document: an exact source match, a partial or multi-block match, a fuzzy match, or not_found when the value couldn’t be located in the recognized content — the strongest signal that a field deserves review.

Source grounding is on by default. When only the extracted values matter, turn it off with settings.IncludeSourceLocations = false to cut model token usage — grounding asks the model to also return per-field source references, which roughly doubles the schema sent with each request.

For per-field confidence signals, set settings.IncludeConfidence = true. Each metadata entry then also carries the individual confidence components for the field, and the run emits the confidence.json sidecar. Combined with the match grounding labels, this gives your pipeline a per-field basis for automatic acceptance versus human review.

Error handling

The SDK throws a NutrientException when extraction fails:

catch (NutrientException e)
{
Console.Error.WriteLine($"Error: {e.Message}");
Environment.Exit(1);
}

Common failure scenarios include:

  • The document can’t be read due to path or permission issues
  • The JSON Schema is missing or malformed (validated before any model call)
  • The model endpoint is unreachable, the model can’t produce a schema-conforming response, or the feature isn’t licensed

In production code:

  • Catch NutrientException.
  • Return a clear error message.
  • Log failure details for debugging.

Conclusion

The workflow for structured data extraction is:

  1. Open the source document, and bind a DataExtraction instance to it with DataExtraction.Set() — both in a using statement(opens in a new tab).
  2. Point ExtractionSchemaFile (or ExtractionSchema) at the JSON Schema to fill, and add optional instructions.
  3. Configure the AI provider on the document settings — local for privacy, or a hosted provider — or inherit it from the SDK-wide settings.
  4. Call Extract() and read the result.json artifact; consume the extraction node and use metadata to verify or review.
  5. Check HasArtifact for the conditional artifacts, or call SaveTo() to write the whole set.
  6. Handle NutrientException for robust error recovery.

For related extraction workflows, refer to the .NET SDK guides.

Download this ready-to-use sample package to explore structured data extraction.