---
title: "Extracting data from images using OCR | Nutrient .NET SDK"
canonical_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/extract-data-from-image-ocr/"
md_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/extract-data-from-image-ocr.md"
last_updated: "2026-10-08T00:00:00.000Z"
description: "Extract text from images using fast OCR with Nutrient .NET SDK. Optimized for high-throughput processing and simple text-based documents."
---

# Extracting data from images using OCR

Use Adaptive OCR to extract text from images for high-throughput workflows.

Common use cases include:

- Invoice and receipt processing

- Search indexing pipelines

- Real-time text capture

- Large-scale document digitization

OCR focuses on text extraction and word-level coordinates. It doesn't perform full semantic layout analysis like ICR.

[Download sample](https://www.nutrient.io/downloads/samples/csharp/extract-data-from-image-ocr.zip)

## How Nutrient helps

Nutrient.NET SDK handles Adaptive OCR configuration, extraction, and JSON output.

The SDK handles:

- OCR engine and model configuration details

- Word-level bounding box calculations

- Text line detection and reading order handling

- Multi-language recognition internals

## Prerequisites

Before following this guide, ensure you have:

-.NET 8.0 or higher installed

- Nutrient.NET SDK referenced by your project

- An image file to process (PNG, JPEG, or other supported formats)

For initial SDK setup and configuration, refer to the [getting started](https://www.nutrient.io/sdk/dotnet/getting-started/csharp.md) guide.

## Complete implementation

This example extracts OCR text and writes the output as JSON:

```csharp

using Nutrient;

```

## Configuring Adaptive OCR mode

Open the image with a [using statement](https://learn.microsoft.com/dotnet/csharp/language-reference/statements/using) and set the vision engine to Adaptive OCR.

In this sample:

- Setting `Engine` to `VisionEngine.AdaptiveOcr` enables Adaptive OCR mode.

- For image inputs like this sample, Adaptive OCR behaves like a fast OCR extraction pipeline.

```csharp

try
{
    using Document document = Document.Open("input_ocr_multiple_languages.png");
    // Configure OCR engine for fast text extraction
    document.Settings.VisionSettings.Engine = VisionEngine.AdaptiveOcr;

```

## Creating a vision instance and extracting content

Create a vision instance and call `ExtractContent()`.

In this sample:

- `Vision.Set(document)` binds OCR extraction to the opened document.

- `ExtractContent()` returns OCR results as a JSON string.

- The output includes extracted text and coordinates.

```csharp

    var vision = Vision.Set(document);
    string contentJson = vision.ExtractContent();

```

Write the JSON string to a file for downstream processing.

Use this output for indexing, analytics, or storage:

```csharp

    File.WriteAllText("output.json", contentJson);
}
catch (NutrientException e)
{
    Console.Error.WriteLine($"Error: {e.Message}");
    Environment.Exit(1);
}

```

## Understanding the output

`ExtractContent()` in Adaptive OCR mode returns JSON optimized for text and word-level positions.

OCR output includes:

- **Text content** — Extracted text with line structure

- **Bounding boxes** — Pixel coordinates for text regions

- **Word-level data** — Per-word positions for highlighting or targeting

- **Language detection** — May be available in OCR output depending on the content and extraction result

### Key output fields

These are the most commonly used fields in OCR JSON output:

- **`text`** — Extracted text for the element.

- **`words`** — Per-word OCR results.

- **`bounds`** — Bounding box coordinates for the element or word.

- **`confidence`** — Confidence score for the element or word.

- **`readingOrder`** — Sequence in which elements should be read.

- **`id`** — Unique identifier for the extracted element.

- **`pageNumber`** — Source page number.

- **`type` / `role`** — Semantic type of the extracted block when available.

When an element contains only one word, element-level and word-level `bounds`/`confidence` can appear identical.

Unlike ICR output, OCR output focuses on text and positions instead of semantic document structure.

## Error handling

Vision API throws a `NutrientException` when OCR extraction fails.

Common failure scenarios include:

- The image file can't be read because of path or permission issues.

- Image data is corrupted or uses unsupported encoding.

- OCR models are missing or inaccessible.

- The available memory is insufficient for large images.

- The image format or resolution is unsupported.

In production code:

- Catch `NutrientException`.

- Return a clear error message.

- Log failure details for debugging.

## Conclusion

Use this workflow for Adaptive OCR-based text extraction:

1. Open the image document with a `using` statement for automatic resource cleanup.

2. Set `Engine` to `VisionEngine.AdaptiveOcr` in the vision settings for fast text extraction.

3. For image inputs, Adaptive OCR focuses on character recognition and word extraction without semantic analysis or layout detection.

4. Create a vision instance with `Vision.Set()` to bind text extraction operations to the document.

5. Call `ExtractContent()` to invoke the OCR engine for character recognition.

6. The OCR engine performs word detection, calculates bounding boxes, and generates JSON output with text and coordinates.

7. The method returns a JSON-formatted string containing extracted text with word-level bounding boxes in pixel coordinates.

8. OCR processing is optimized for speed, minimizing computational overhead for high-throughput scenarios.

9. Write the JSON content to a file for search indexing (Elasticsearch, Solr), text analysis, and database storage.

10. Handle `NutrientException` failures for robust error recovery in production environments.

11. Adaptive OCR mode is ideal for invoice processing, receipt scanning, search indexing, and document digitization where speed is critical.

For related image extraction workflows, refer to the [.NET SDK guides](https://www.nutrient.io/guides/dotnet/csharp.md).

Download [this ready-to-use sample package](https://www.nutrient.io/downloads/samples/csharp/extract-data-from-image-ocr.zip) to explore the Vision API capabilities with preconfigured OCR settings.
---

## Related pages

- [Nutrient .NET SDK extraction guides](/guides/dotnet/csharp/extraction.md)
- [Applying OCR to a PDF page](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf-page.md)
- [Applying OCR to a PDF document](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf.md)
- [Classifying documents](/guides/dotnet/csharp/extraction/classify-document.md)
- [Generating image descriptions using Claude](/guides/dotnet/csharp/extraction/describe-image-with-claude.md)
- [Generating image descriptions using local AI](/guides/dotnet/csharp/extraction/describe-image-with-local-ai.md)
- [Generating image descriptions using OpenAI](/guides/dotnet/csharp/extraction/describe-image-with-openai.md)
- [Detecting document language](/guides/dotnet/csharp/extraction/detect-document-language.md)
- [Extracting data from images using ICR](/guides/dotnet/csharp/extraction/extract-data-from-image-icr.md)
- [Extracting data from images using vision language models](/guides/dotnet/csharp/extraction/extract-data-from-image-vlm.md)
- [Extracting data from specific pages](/guides/dotnet/csharp/extraction/extract-data-from-specific-pages.md)
- [Extracting form fields from images](/guides/dotnet/csharp/extraction/extract-form-fields-from-image.md)
- [Extracting structured data from documents](/guides/dotnet/csharp/extraction/extract-structured-data.md)
- [Generating extraction schemas](/guides/dotnet/csharp/extraction/generate-extraction-schema.md)
- [Extracting JSON data from a PDF document](/guides/dotnet/csharp/extraction/json-data-extraction.md)
- [Labeling form fields with a vision language model](/guides/dotnet/csharp/extraction/label-form-fields-with-vlm.md)
- [Parsing a document into structured content](/guides/dotnet/csharp/extraction/parse-document.md)
- [Extracting text from PDF documents](/guides/dotnet/csharp/extraction/pdf-to-text.md)
- [Reading barcodes with vision extraction](/guides/dotnet/csharp/extraction/read-barcodes-with-vision.md)
- [Extracting text from multilingual images](/guides/dotnet/csharp/extraction/read-text-from-image-multi-language.md)
- [Extracting text from images](/guides/dotnet/csharp/extraction/read-text-from-image.md)
- [Searching document text](/guides/dotnet/csharp/extraction/search-document-text.md)
- [Speeding up first ICR operation by predownloading models](/guides/dotnet/csharp/extraction/speed-up-first-icr-by-downloading-requirements.md)

