Extracting structured data from documents
Most document workflows don’t want a wall of recognized text — they want fields: the invoice total, the patient’s date of birth, every line item as a row. Structured extraction turns a document into exactly the JSON you ask for: you supply a JSON Schema describing the fields, and an AI model fills it from the document’s recognized content.
This sample shows how to extract schema-shaped data from a document using Nutrient .NET SDK. The result reports not just the values but also where each value came from — per-field source locations and grounding labels you can use to verify the extraction against the original document.
Download sampleHow Nutrient helps
The DataExtraction class runs the full structured extraction workflow behind a single method call. It runs the same extract operation, under the same option names, and returns the same output files as the hosted Nutrient data extraction service, so the code and the option names you learn here carry over between the two. The SDK handles:
- Reading the document and recognizing text, tables, key-value regions, and form fields in reading order
- Sending the recognized content and your JSON Schema to the AI model as a structured-output request
- Retrying automatically when the model’s response doesn’t conform to the schema
- Grounding each extracted value back to its source location in the document
- Serializing the result to JSON
A successful extraction conforms to your schema — the same call with the same schema yields the same shape, ready for your downstream code to consume without defensive parsing. If the model exhausts its retries, the default behavior is to return an empty extraction with a warning. The sample instead turns that fallback off so a failed extraction throws rather than looking like valid business data.
Parse isn’t a prerequisite
Extract and Parse are two independent operations on the same class. Extract reads the document itself: it runs the extraction pipeline internally, feeds the recognized content to the model, and returns the filled schema. You don’t have to call Parse first, and passing parsed output into Extract isn’t a supported input. Call Parse when you want the document’s structure as JSON or Markdown; call Extract when you want your schema’s fields filled in.
What shapes the result
Two inputs shape the result:
- JSON Schema (required) — The fields to extract. Supply it inline with
ExtractionSchemaor from a file withExtractionSchemaFile; the two are mutually exclusive. Each schema property’sdescriptiontells the extractor what belongs there — the better the description, the better the match. - Instructions (optional) — Free-form guidance for the extraction: disambiguation rules, formatting preferences, or domain context — anything you’d tell a colleague doing the extraction by hand.
Extraction requires an AI model. Configure the provider on the document’s ExtractSettings; it’s the same value as AiProcessingSettings, so a provider set for the whole SDK applies too. A local OpenAI-compatible server keeps documents on your machine; a hosted provider needs an API key. Structured extraction requires the vision data extraction feature in your license.
Complete implementation
Import the Nutrient namespace:
using Nutrient;Opening the document
Open the source document and bind a DataExtraction instance to it. Both are disposable, so use C#’s using statement(opens in a new tab) for each — the document releases the file, and the DataExtraction instance drops its reference so the native bridge can free the handle that pins the document:
try{ using Document document = Document.Open("input_invoice_lumen.pdf"); using DataExtraction dataExtraction = DataExtraction.Set(document);Supplying the schema
Like every SDK operation, Extract() reads its options from the document settings. ExtractSettings holds the options of the hosted extraction service under the same names; options you don’t set keep the default the extraction pipeline or the SDK-wide settings provide.
This sample ships input_invoice_lumen_schema.json, a JSON Schema describing the invoice fields to pull out. Point ExtractionSchemaFile at it rather than embedding the schema in your source, so the schema stays editable without a rebuild.
The settings wrapper holds a native handle, so declare it with using as well. Ask for a failed extraction to throw instead of returning an empty result:
using ExtractSettings settings = document.Settings.ExtractSettings; settings.ExtractionSchemaFile = "input_invoice_lumen_schema.json"; settings.Instructions = "Amounts are plain numbers without currency symbols."; document.Settings.AiProcessingSettings.UnparseableOutputHandling = UnparseableOutputHandling.Fail;The bundled schema is a plain JSON Schema object — type, properties, and required — where every property carries the description the extractor matches against the document:
{ "type": "object", "properties": { "invoice number": { "type": "string", "description": "The unique invoice number or identifier" }, "total": { "type": "number", "description": "The final total amount of the invoice including all taxes and discounts." } }, "required": ["invoice number", "total"]}To build the schema at runtime instead, assign the same JSON to settings.ExtractionSchema as a string.
Configuring the AI provider
Point the run at your model. This example uses a local OpenAI-compatible endpoint, so the document never leaves your machine:
settings.Provider = "local"; settings.Endpoint = "http://localhost:1234/v1"; settings.Model = "your-model-id";For a hosted provider instead, set Provider to "openai" or "anthropic" along with ApiKey (and Endpoint for an OpenAI-compatible proxy).
Strict structured output
Strict structured output is enabled by default through the document settings. The sample assigns it explicitly so the behavior is clear: the model’s response is grammar-constrained to the schema, and a successful result matches it exactly.
settings.StrictStructuredOutput = true;Two things to know about strict mode:
- You don’t change your schema. Strict mode has formal requirements (every object closed, every property accounted for), and the SDK normalizes your schema to satisfy them automatically.
- Absent fields come back as
null. In strict mode the model must emit every schema property, so a field the document doesn’t contain is returned asnullinstead of being omitted — your downstream code can rely on every key being present. Fields you list in the schema’srequiredarray keep their declared type untouched — they can only benullif your own schema allows it.
Strict mode requires a model and endpoint that support grammar-constrained structured outputs. Check your provider or local server’s documentation. If it doesn’t support them, set StrictStructuredOutput to false; the SDK then validates the model response and retries when it doesn’t conform.
Extracting the data
Call Extract() and read the result.json artifact. Every run produces it, so it’s safe to read by name. The result wrapper holds native resources as well, so it gets a using declaration too — read what you need from it before it goes out of scope at the end of the block:
using ExtractResult result = dataExtraction.Extract(); File.WriteAllText("output.json", result.GetArtifact("result.json")); Console.WriteLine("Extraction written to output.json");Handling the conditional artifacts
Besides result.json, a run can emit up to three more artifacts, each only when the conditions for it are met. Check with HasArtifact before reading, because GetArtifact throws on a name the run didn’t produce:
if (result.HasArtifact("warnings.json")) { Console.WriteLine($"Warnings: {result.GetArtifact("warnings.json")}"); }
Console.WriteLine($"Artifacts: {string.Join(", ", result.ArtifactNames)}");}| Artifact | Emitted when |
|---|---|
result.json | Always. |
warnings.json | At least one non-fatal AI-layer warning was collected. |
confidence.json | IncludeConfidence is true and at least one scoreable signal exists. |
diagnostics.json | IncludeDiagnostics is true. |
result.SaveTo("artifacts") writes every artifact the run produced into a directory, each under its own name — useful when you want the whole set on disk without checking each one.
Understanding the output
result.json has two top-level nodes:
extraction— The extracted fields, shaped exactly to your schema.metadata— One entry per extracted field, carrying where the value came from: amatchgrounding label and source location info (page and bounding box) so you can highlight the source in a viewer or route low-trust fields to human review.
The match label tells you how confidently the value was traced back to the document: an exact source match, a partial or multi-block match, a fuzzy match, or not_found when the value couldn’t be located in the recognized content — the strongest signal that a field deserves review.
Source grounding is on by default. When only the extracted values matter, turn it off with settings.IncludeSourceLocations = false to cut model token usage — grounding asks the model to also return per-field source references, which roughly doubles the schema sent with each request.
For per-field confidence signals, set settings.IncludeConfidence = true. Each metadata entry then also carries the individual confidence components for the field, and the run emits the confidence.json sidecar. Combined with the match grounding labels, this gives your pipeline a per-field basis for automatic acceptance versus human review.
Error handling
The SDK throws a NutrientException when extraction fails:
catch (NutrientException e){ Console.Error.WriteLine($"Error: {e.Message}"); Environment.Exit(1);}Common failure scenarios include:
- The document can’t be read due to path or permission issues
- The JSON Schema is missing or malformed (validated before any model call)
- The model endpoint is unreachable, the model can’t produce a schema-conforming response, or the feature isn’t licensed
In production code:
- Catch
NutrientException. - Return a clear error message.
- Log failure details for debugging.
Conclusion
The workflow for structured data extraction is:
- Open the source document, and bind a
DataExtractioninstance to it withDataExtraction.Set()— both in a using statement(opens in a new tab). - Point
ExtractionSchemaFile(orExtractionSchema) at the JSON Schema to fill, and add optional instructions. - Configure the AI provider on the document settings — local for privacy, or a hosted provider — or inherit it from the SDK-wide settings.
- Call
Extract()and read theresult.jsonartifact; consume theextractionnode and usemetadatato verify or review. - Check
HasArtifactfor the conditional artifacts, or callSaveTo()to write the whole set. - Handle
NutrientExceptionfor robust error recovery.
For related extraction workflows, refer to the .NET SDK guides.
Download this ready-to-use sample package to explore structured data extraction.