---
title: "Parse documents in C# | Nutrient .NET SDK"
canonical_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/parse-document/"
md_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/parse-document.md"
last_updated: "2026-10-08T00:00:00.000Z"
description: "Parsing a PDF or image into structured JSON and Markdown content using Nutrient .NET SDK."
---

# Parsing a document into structured content

Document parsing turns a PDF or image into reusable structured content. Use it when you need paragraphs, tables, figures, key-value regions, words, or Markdown rather than a converted document or a fixed set of extracted fields.

This guide uses `DataExtraction.Parse` to produce JSON and Markdown artifacts from the same document. The operation uses the same option names and output filenames as the hosted Nutrient data extraction service.

[Download sample](https://www.nutrient.io/downloads/samples/csharp/parse-document.zip)

## Preparing the project

Import the Nutrient namespace:

```csharp

using Nutrient;

```

## Opening the document

Open the source document and bind a `DataExtraction` instance to it. Both are disposable, so use C#’s [using statement](https://learn.microsoft.com/dotnet/csharp/language-reference/statements/using) for each:

```csharp

try
{
    using Document document = Document.Open("input_invoice_lumen.pdf");
    using DataExtraction dataExtraction = DataExtraction.Set(document);

```

The sample uses an invoice, but Parse works with any supported PDF, TIFF, or image.

## Selecting the output formats

Like every SDK operation, `Parse()` reads its options from the document settings. `ParseSettings` holds the options of the hosted parse service under the same names; options you don't set keep the pipeline or SDK-wide default.

Set `ExportFormats` to request structured JSON and Markdown in one pass. Both the settings wrapper and the result wrapper hold native resources, so declare each with `using`:

```csharp

    using ParseSettings settings = document.Settings.ParseSettings;
    settings.ExportFormats = new[] { "json", "markdown" };
    settings.IncludeWords = true;

    using ParseResult result = dataExtraction.Parse();

```

`IncludeWords` adds word-level details and bounding boxes to the JSON output. Leave it unassigned when the higher-level document structure is enough.

## Reading named artifacts

Parse returns named artifacts instead of one output string. A run requesting JSON and Markdown produces `document.json` and `document.md`:

```csharp

    Console.WriteLine($"Artifacts: {string.Join(", ", result.ArtifactNames)}");

    File.WriteAllText("document.json", result.GetArtifact("document.json"));
    File.WriteAllText("document.md", result.GetArtifact("document.md"));
    Console.WriteLine("Parsed content written to document.json and document.md");
}
catch (NutrientException e)
{
    Console.Error.WriteLine($"Error: {e.Message}");
    Environment.Exit(1);
}

```

`GetArtifact` throws when the requested artifact isn’t present. If formats are selected dynamically, call `HasArtifact` before reading. The typed accessors `DocumentJson` and `DocumentMd` return the same content, or `null` when the run didn’t produce that artifact.

Call `result.SaveTo("artifacts")` when you want to write every produced artifact into one directory under its service-compatible filename.

## Choosing between Parse, conversion, and Extract

Use the operation that matches the output you need:

- Use `Document.ExportAsMarkdown` when you only need a direct PDF-to-Markdown conversion.

- Use `DataExtraction.Parse` when you need reusable document structure, several output formats, word details, or service-compatible artifacts.

- Use `DataExtraction.Extract` when you want an AI model to fill fields defined by your JSON Schema. Calling Parse first isn’t required.

## Conclusion

The parsing workflow is:

1. Open the document and bind `DataExtraction` with `DataExtraction.Set()`.

2. Select JSON, Markdown, or both with the document's `ParseSettings`.

3. Call `Parse()` once.

4. Read the named artifacts or save the complete result.

5. Handle `NutrientException` for failures.

Download [this ready-to-use sample package](https://www.nutrient.io/downloads/samples/csharp/parse-document.zip) to explore document parsing.
---

## Related pages

- [Nutrient .NET SDK extraction guides](/guides/dotnet/csharp/extraction.md)
- [Applying OCR to a PDF page](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf-page.md)
- [Applying OCR to a PDF document](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf.md)
- [Classifying documents](/guides/dotnet/csharp/extraction/classify-document.md)
- [Generating image descriptions using Claude](/guides/dotnet/csharp/extraction/describe-image-with-claude.md)
- [Generating image descriptions using local AI](/guides/dotnet/csharp/extraction/describe-image-with-local-ai.md)
- [Generating image descriptions using OpenAI](/guides/dotnet/csharp/extraction/describe-image-with-openai.md)
- [Detecting document language](/guides/dotnet/csharp/extraction/detect-document-language.md)
- [Extracting data from images using ICR](/guides/dotnet/csharp/extraction/extract-data-from-image-icr.md)
- [Extracting data from images using OCR](/guides/dotnet/csharp/extraction/extract-data-from-image-ocr.md)
- [Extracting data from images using vision language models](/guides/dotnet/csharp/extraction/extract-data-from-image-vlm.md)
- [Extracting data from specific pages](/guides/dotnet/csharp/extraction/extract-data-from-specific-pages.md)
- [Extracting form fields from images](/guides/dotnet/csharp/extraction/extract-form-fields-from-image.md)
- [Extracting structured data from documents](/guides/dotnet/csharp/extraction/extract-structured-data.md)
- [Generating extraction schemas](/guides/dotnet/csharp/extraction/generate-extraction-schema.md)
- [Extracting JSON data from a PDF document](/guides/dotnet/csharp/extraction/json-data-extraction.md)
- [Labeling form fields with a vision language model](/guides/dotnet/csharp/extraction/label-form-fields-with-vlm.md)
- [Extracting text from PDF documents](/guides/dotnet/csharp/extraction/pdf-to-text.md)
- [Reading barcodes with vision extraction](/guides/dotnet/csharp/extraction/read-barcodes-with-vision.md)
- [Extracting text from multilingual images](/guides/dotnet/csharp/extraction/read-text-from-image-multi-language.md)
- [Extracting text from images](/guides/dotnet/csharp/extraction/read-text-from-image.md)
- [Searching document text](/guides/dotnet/csharp/extraction/search-document-text.md)
- [Speeding up first ICR operation by predownloading models](/guides/dotnet/csharp/extraction/speed-up-first-icr-by-downloading-requirements.md)

