Parsing a document into structured content
Document parsing turns a PDF or image into reusable structured content. Use it when you need paragraphs, tables, figures, key-value regions, words, or Markdown rather than a converted document or a fixed set of extracted fields.
This guide uses DataExtraction.Parse to produce JSON and Markdown artifacts from the same document. The operation uses the same option names and output filenames as the hosted Nutrient data extraction service.
Preparing the project
Import the Nutrient namespace:
using Nutrient;Opening the document
Open the source document and bind a DataExtraction instance to it. Both are disposable, so use C#’s using statement(opens in a new tab) for each:
try{ using Document document = Document.Open("input_invoice_lumen.pdf"); using DataExtraction dataExtraction = DataExtraction.Set(document);The sample uses an invoice, but Parse works with any supported PDF, TIFF, or image.
Selecting the output formats
Like every SDK operation, Parse() reads its options from the document settings. ParseSettings holds the options of the hosted parse service under the same names; options you don’t set keep the pipeline or SDK-wide default.
Set ExportFormats to request structured JSON and Markdown in one pass. Both the settings wrapper and the result wrapper hold native resources, so declare each with using:
using ParseSettings settings = document.Settings.ParseSettings; settings.ExportFormats = new[] { "json", "markdown" }; settings.IncludeWords = true;
using ParseResult result = dataExtraction.Parse();IncludeWords adds word-level details and bounding boxes to the JSON output. Leave it unassigned when the higher-level document structure is enough.
Reading named artifacts
Parse returns named artifacts instead of one output string. A run requesting JSON and Markdown produces document.json and document.md:
Console.WriteLine($"Artifacts: {string.Join(", ", result.ArtifactNames)}");
File.WriteAllText("document.json", result.GetArtifact("document.json")); File.WriteAllText("document.md", result.GetArtifact("document.md")); Console.WriteLine("Parsed content written to document.json and document.md");}catch (NutrientException e){ Console.Error.WriteLine($"Error: {e.Message}"); Environment.Exit(1);}GetArtifact throws when the requested artifact isn’t present. If formats are selected dynamically, call HasArtifact before reading. The typed accessors DocumentJson and DocumentMd return the same content, or null when the run didn’t produce that artifact.
Call result.SaveTo("artifacts") when you want to write every produced artifact into one directory under its service-compatible filename.
Choosing between Parse, conversion, and Extract
Use the operation that matches the output you need:
- Use
Document.ExportAsMarkdownwhen you only need a direct PDF-to-Markdown conversion. - Use
DataExtraction.Parsewhen you need reusable document structure, several output formats, word details, or service-compatible artifacts. - Use
DataExtraction.Extractwhen you want an AI model to fill fields defined by your JSON Schema. Calling Parse first isn’t required.
Conclusion
The parsing workflow is:
- Open the document and bind
DataExtractionwithDataExtraction.Set(). - Select JSON, Markdown, or both with the document’s
ParseSettings. - Call
Parse()once. - Read the named artifacts or save the complete result.
- Handle
NutrientExceptionfor failures.
Download this ready-to-use sample package to explore document parsing.