---
title: "Applying OCR to a PDF document | Nutrient .NET SDK"
canonical_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/apply-ocr-to-pdf/"
md_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/apply-ocr-to-pdf.md"
last_updated: "2026-10-08T00:00:00.000Z"
description: "How to run OCR on a PDF document using Nutrient .NET SDK."
---

# Applying OCR to a PDF document

Scanned documents arrive in many formats: image-based PDFs, multi-page TIFFs, single PNG or JPG scans, or even faxed exports. In every case the pages are pictures of text rather than selectable, searchable content. Legal firms receiving historical case files, healthcare teams handling medical records, and finance departments archiving older statements all share this problem: the information is visible but not searchable.

Applying OCR adds an invisible text layer that sits behind the page image. The visual appearance doesn't change, but the text becomes selectable, searchable, and readable by assistive technologies.

This sample shows how to run OCR over every page of a document using Nutrient.NET SDK and produce a searchable PDF as output. The input can be any document format the SDK supports. If the input isn't already a PDF, the SDK converts it to PDF automatically when you create the editor.

[Download sample](https://www.nutrient.io/downloads/samples/csharp/apply-ocr-to-pdf.zip)

## How Nutrient helps

Nutrient.NET SDK handles the full OCR pipeline behind a single method call. The SDK takes care of:

- Implicitly converting non-PDF inputs (images, multi-page TIFFs, Office documents) to PDF when the editor is created

- Rendering each PDF page to a bitmap at the resolution OCR needs

- Running text recognition with the configured languages

- Preserving reading order and text block orientation returned by the recognizer

- Placing an invisible, correctly positioned text layer over the original page content

You control the outcome through document settings, such as which OCR languages to use.

## Preparing the project

Import the Nutrient namespace:

```csharp

using Nutrient;

```

## Running OCR on the whole document

Open the source document inside a [using statement](https://learn.microsoft.com/dotnet/csharp/language-reference/statements/using), configure the OCR language, and call `MakeSearchable()` on the editor. The sample passes an image-based PDF as input, but the same code handles raw images, multi-page TIFFs, or any other supported document format. The `using` statements close the document and editor automatically when the block ends, even if an error is thrown:

```csharp

try
{
    using Document document = Document.Open("input_image_based.pdf");
    document.Settings.OcrSettings.DefaultLanguages = "eng";

    using PdfEditor editor = PdfEditor.Edit(document);
    editor.MakeSearchable();

```

Setting `DefaultLanguages` to `"eng"` tells the recognizer which language models to load. Combine languages with `+` (for example `"eng+deu"`) when you know the document contains more than one language. Setting it to match the document content directly improves accuracy on ambiguous characters.

`PdfEditor.Edit(document)` attaches an editor to the open document. If the document isn't already a PDF, the SDK converts it to PDF at this step so the rest of the pipeline works on a uniform page representation. Calling `editor.MakeSearchable()` loops through every page, runs OCR, and writes an invisible text layer on top of the existing page content. Any hidden text already present on a page is removed before the new layer is drawn, so re-running OCR doesn't duplicate content.

## Saving the result

Save the modified document to a new file. Wrap the workflow in `try/catch` on `NutrientException` to surface any licensing, language-pack, or I/O issue that the SDK reports:

```csharp

    editor.SaveAs("output.pdf");
    Console.WriteLine("Successfully applied OCR to output.pdf");
}
catch (NutrientException e)
{
    Console.Error.WriteLine($"Error: {e.Message}");
    Environment.Exit(1);
}

```

## Conclusion

The workflow for OCR-ing a whole PDF is:

1. Open the source document.

2. Configure OCR languages on the document settings.

3. Create a `PdfEditor` for the document.

4. Call `MakeSearchable()` to apply OCR to every page.

5. Save the result — the `using` statements release the editor and document.

The output is a standard PDF with an invisible text layer, so existing PDF viewers, search tools, and accessibility software can read it without any extra configuration.
---

## Related pages

- [Nutrient .NET SDK extraction guides](/guides/dotnet/csharp/extraction.md)
- [Applying OCR to a PDF page](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf-page.md)
- [Classifying documents](/guides/dotnet/csharp/extraction/classify-document.md)
- [Generating image descriptions using Claude](/guides/dotnet/csharp/extraction/describe-image-with-claude.md)
- [Generating image descriptions using local AI](/guides/dotnet/csharp/extraction/describe-image-with-local-ai.md)
- [Generating image descriptions using OpenAI](/guides/dotnet/csharp/extraction/describe-image-with-openai.md)
- [Detecting document language](/guides/dotnet/csharp/extraction/detect-document-language.md)
- [Extracting data from images using ICR](/guides/dotnet/csharp/extraction/extract-data-from-image-icr.md)
- [Extracting data from images using OCR](/guides/dotnet/csharp/extraction/extract-data-from-image-ocr.md)
- [Extracting data from images using vision language models](/guides/dotnet/csharp/extraction/extract-data-from-image-vlm.md)
- [Extracting data from specific pages](/guides/dotnet/csharp/extraction/extract-data-from-specific-pages.md)
- [Extracting form fields from images](/guides/dotnet/csharp/extraction/extract-form-fields-from-image.md)
- [Extracting structured data from documents](/guides/dotnet/csharp/extraction/extract-structured-data.md)
- [Generating extraction schemas](/guides/dotnet/csharp/extraction/generate-extraction-schema.md)
- [Extracting JSON data from a PDF document](/guides/dotnet/csharp/extraction/json-data-extraction.md)
- [Labeling form fields with a vision language model](/guides/dotnet/csharp/extraction/label-form-fields-with-vlm.md)
- [Parsing a document into structured content](/guides/dotnet/csharp/extraction/parse-document.md)
- [Extracting text from PDF documents](/guides/dotnet/csharp/extraction/pdf-to-text.md)
- [Reading barcodes with vision extraction](/guides/dotnet/csharp/extraction/read-barcodes-with-vision.md)
- [Extracting text from multilingual images](/guides/dotnet/csharp/extraction/read-text-from-image-multi-language.md)
- [Extracting text from images](/guides/dotnet/csharp/extraction/read-text-from-image.md)
- [Searching document text](/guides/dotnet/csharp/extraction/search-document-text.md)
- [Speeding up first ICR operation by predownloading models](/guides/dotnet/csharp/extraction/speed-up-first-icr-by-downloading-requirements.md)

