---
title: "Extracting text from PDF documents | Nutrient .NET SDK"
canonical_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/pdf-to-text/"
md_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/pdf-to-text.md"
last_updated: "2026-10-08T00:00:00.000Z"
description: "Extract layout-preserving text from PDF documents using Nutrient .NET SDK."
---

# Extracting text from PDF documents

PDF-to-text extraction pulls readable content from a static document while preserving its spatial arrangement. Layout-aware extraction keeps columns, indentation, and table alignment intact, so the output matches what readers see on the page.

Use programmatic extraction to:

- Index large document libraries for search.

- Send structured text to data pipelines and language models.

- Reuse report and statement content without manual retyping.

## Extract PDF text with the.NET SDK

You can add layout-preserving text extraction to a.NET application with the Nutrient.NET SDK. The SDK extracts text directly from PDFs, so you don't need external tools for this workflow.

## Prepare the project

Start by importing the Nutrient namespace:

```csharp

using Nutrient;

```

## Load the PDF document

This guide uses the `Document` class. Use C#'s [using statement](https://learn.microsoft.com/dotnet/csharp/language-reference/statements/using) to manage the document instance lifecycle.

The SDK can load a source file from a file path or a stream. This guide uses a file path:

```csharp

try
{
    using Document document = Document.Open("input.pdf");

```

The path can be absolute or relative. This example loads the file from the application's working directory.

## Extract layout-preserving text

Call `ExportAsText` to extract the document text into a plain-text file. The method maps each word to a character grid that mirrors its position on the page:

```csharp

    document.ExportAsText("output.txt");
    Console.WriteLine("Successfully extracted to output.txt");
}
catch (NutrientException e)
{
    Console.Error.WriteLine($"Error: {e.Message}");
    Environment.Exit(1);
}

```

The `ExportAsText` method analyzes the PDF text content and the position of each word, then reconstructs the page in plain text. Words that sit close together join with single spaces, large horizontal gaps become proportional whitespace that preserves columns and tab stops, and vertical gaps between lines produce blank lines. The result reads like the original page while staying in a portable format.

The method handles these PDF content types:

- Flowing text.

- Multi-column layouts.

- Tables and aligned data.

- Mixed content layouts.

## Handle errors

Nutrient.NET SDK uses exception handling for errors. The methods in this guide throw a `NutrientException` if a failure occurs. Use this exception to troubleshoot issues and implement error handling logic.

## Conclusion

You've extracted layout-preserving text from a PDF document. The extracted content is ready for search indexing, data pipelines, and downstream processing.
---

## Related pages

- [Nutrient .NET SDK extraction guides](/guides/dotnet/csharp/extraction.md)
- [Applying OCR to a PDF page](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf-page.md)
- [Applying OCR to a PDF document](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf.md)
- [Classifying documents](/guides/dotnet/csharp/extraction/classify-document.md)
- [Generating image descriptions using Claude](/guides/dotnet/csharp/extraction/describe-image-with-claude.md)
- [Generating image descriptions using local AI](/guides/dotnet/csharp/extraction/describe-image-with-local-ai.md)
- [Generating image descriptions using OpenAI](/guides/dotnet/csharp/extraction/describe-image-with-openai.md)
- [Detecting document language](/guides/dotnet/csharp/extraction/detect-document-language.md)
- [Extracting data from images using ICR](/guides/dotnet/csharp/extraction/extract-data-from-image-icr.md)
- [Extracting data from images using OCR](/guides/dotnet/csharp/extraction/extract-data-from-image-ocr.md)
- [Extracting data from images using vision language models](/guides/dotnet/csharp/extraction/extract-data-from-image-vlm.md)
- [Extracting data from specific pages](/guides/dotnet/csharp/extraction/extract-data-from-specific-pages.md)
- [Extracting form fields from images](/guides/dotnet/csharp/extraction/extract-form-fields-from-image.md)
- [Extracting structured data from documents](/guides/dotnet/csharp/extraction/extract-structured-data.md)
- [Generating extraction schemas](/guides/dotnet/csharp/extraction/generate-extraction-schema.md)
- [Extracting JSON data from a PDF document](/guides/dotnet/csharp/extraction/json-data-extraction.md)
- [Labeling form fields with a vision language model](/guides/dotnet/csharp/extraction/label-form-fields-with-vlm.md)
- [Parsing a document into structured content](/guides/dotnet/csharp/extraction/parse-document.md)
- [Reading barcodes with vision extraction](/guides/dotnet/csharp/extraction/read-barcodes-with-vision.md)
- [Extracting text from multilingual images](/guides/dotnet/csharp/extraction/read-text-from-image-multi-language.md)
- [Extracting text from images](/guides/dotnet/csharp/extraction/read-text-from-image.md)
- [Searching document text](/guides/dotnet/csharp/extraction/search-document-text.md)
- [Speeding up first ICR operation by predownloading models](/guides/dotnet/csharp/extraction/speed-up-first-icr-by-downloading-requirements.md)

