This HTML page is not optimized for LLM or AI agent consumption. Fetch the Markdown version instead: /guides/dotnet/csharp/extraction/json-data-extraction.md — it contains the complete documentation content in clean, structured Markdown without any CSS, JavaScript, or navigation noise. Extracting JSON data from a PDF document | Nutrient .NET SDK

Extract structured data from PDF files as JSON for storage, API workflows, or analytics pipelines. This approach reduces manual entry and gives your application direct access to document content.

Download sample

How Nutrient supports this workflow

Nutrient .NET SDK handles structured extraction from PDF documents, including digital-native PDFs and PDFs that mix digital text with scanned content.

In this sample, VisionEngine.AdaptiveOcr uses an adaptive extraction pipeline that prefers native PDF text when available and falls back to OCR for image-based content when needed.

You don’t need to manage:

  • Third-party OCR engine integration
  • Switching between native-text extraction and OCR
  • Document layout parsing
  • Model download and initialization
  • Conversion from extracted output to structured data

Use the SDK API to extract structured JSON in your application.

Complete implementation

This example shows a complete PDF-to-JSON extraction flow.

Import the Nutrient namespace:

using Nutrient;

Open the PDF document inside a using statement(opens in a new tab). The using statement closes the document automatically:

try
{
using Document document = Document.Open("input.pdf");

Configure the Adaptive OCR engine, extract JSON content, and write it to output.json. Catch NutrientException to handle SDK errors:

document.Settings.VisionSettings.Engine = VisionEngine.AdaptiveOcr;
Vision vision = Vision.Set(document);
string contentJson = vision.ExtractContent();
File.WriteAllText("output.json", contentJson);
Console.WriteLine("Successfully extracted content to output.json");
}
catch (NutrientException e)
{
Console.Error.WriteLine($"Error: {e.Message}");
Environment.Exit(1);
}

Summary

The extraction flow has four steps:

  1. Open the PDF document.
  2. Configure the Adaptive OCR engine.
  3. Extract content as JSON with Vision.
  4. Write the JSON output to a file.

Nutrient handles adaptive extraction and content structuring, so you don’t need to implement PDF parsing, native-text detection, or OCR fallback logic.

Barcode data in JSON output

For documents that contain machine-readable codes, Vision extraction includes detected barcode data in the document layout output. Each detected barcode is represented as a layout element with the decoded value and barcode symbology, such as 1D barcodes, QR codes, Micro QR codes, PDF417, DataMatrix, Aztec, or MaxiCode.

Use this output when a pipeline needs both document text and embedded barcode values from the same pass. To focus specifically on barcodes, refer to the read barcodes with Vision guide.

You can download this sample package to run the example locally.