---
title: "Describe images with local AI in C# | Nutrient .NET SDK"
canonical_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/describe-image-with-local-ai/"
md_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/describe-image-with-local-ai.md"
last_updated: "2026-10-08T00:00:00.000Z"
description: "Generate accessible image descriptions using local AI models with Nutrient .NET SDK."
---

# Generating image descriptions using local AI

Use local AI image description when you need privacy, cost control, or offline operation.

Common use cases include:

- On-premises accessibility workflows

- Medical or regulated environments with local-only data handling

- Secure networks where external APIs are restricted

- High-volume processing without per-image API fees

- Offline-capable applications

This guide uses a local OpenAI-compatible VLM endpoint with Nutrient Vision API.

[Download sample](https://www.nutrient.io/downloads/samples/csharp/describe-image-with-local-ai.zip)

## How Nutrient helps

Nutrient.NET SDK handles local VLM integration, endpoint configuration, and request/response processing.

The SDK handles:

- OpenAI-compatible endpoint formatting and communication

- Image encoding and multimodal payload construction

- Model parameters such as temperature and max tokens

- Local server and model runtime failure handling

## Complete implementation

This example generates image descriptions using a local AI endpoint:

```csharp

using Nutrient;

```

## Opening the image file and configuring the local server

Open the image with a [using statement](https://learn.microsoft.com/dotnet/csharp/language-reference/statements/using) and configure local endpoint settings if needed.

In this sample:

- The SDK supports PNG, JPEG, GIF, BMP, and TIFF.

- The default endpoint is `http://localhost:1234/v1`.

- The default model is `qwen/qwen3-vl-4b`.

- The custom VLM API settings override defaults.

```csharp

try
{
    using Document document = Document.Open("input_photo.png");

    // Optional: Configure the VLM API endpoint settings
    // These settings are customizable based on your VLM provider
    var vlmSettings = document.Settings.CustomVlmApiSettings;
    vlmSettings.ApiEndpoint = "http://localhost:1234/v1";
    vlmSettings.Model = "qwen/qwen3-vl-4b";

```

## Creating a vision instance

Create a vision instance with `Vision.Set(document)`.

Before calling vision methods, ensure:

- The local server is running.

- A vision-capable model is loaded.

- The configured endpoint is reachable.

```csharp

    var vision = Vision.Set(document);

```

## Generating the description

Call `vision.Describe()` to generate a natural language description.

The SDK handles image encoding, request construction, and response parsing:

```csharp

    string description = vision.Describe();

```

## Outputting the description

Print the description for review, or store it in your application.

Common destinations include:

- Database fields

- JSON output files

- HTML `alt` attributes

```csharp

    Console.WriteLine("Image description:");
    Console.WriteLine(description);
}
catch (NutrientException e)
{
    Console.Error.WriteLine($"Error: {e.Message}");
    Environment.Exit(1);
}

```

## Understanding the output

`Describe()` returns natural language text generated by the local model.

Descriptions are typically:

- **Concise** — Focused on key subjects and details, often one to three sentences

- **Accessible** — Suitable for users who rely on screen readers

- **Accurate** — Based on visible content only

- **Model-dependent quality** — Quality varies by model architecture and size

Use this output for accessibility metadata, image search, and document workflows, all with local processing.

## Configuring the VLM API endpoint

Vision API uses OpenAI-compatible endpoints. Configure the custom VLM API settings if your setup differs from defaults:

- `ApiEndpoint` — Base URL for the OpenAI-compatible API (default: `http://localhost:1234/v1`)

- `ApiKey` — API key for secured deployments (optional on many local servers)

- `Model` — Model identifier (default: `qwen/qwen3-vl-4b`)

- `Temperature` — Creativity control (`0.0` deterministic, `1.0` varied phrasing)

- `MaxTokens` — Response token limit (`-1` unlimited)

**Local server setup examples**:

- **LM Studio** — Start server with vision model loaded, default endpoint `http://localhost:1234/v1`

- **Ollama** — Run `ollama serve` with vision model, default endpoint `http://localhost:11434/v1`

- **vLLM** — Launch with `--api-key` flag for authentication, custom port configuration

The SDK handles request formatting and response parsing for local endpoints.

## Error handling

The SDK throws a `NutrientException` when vision operations fail.

Common failure scenarios include:

- The input image can't be read due to path, permission, or format issues

- The local VLM server isn't running or reachable

- No compatible vision model is loaded

- Responses time out on large images or heavy models

- Available CPU/GPU memory is insufficient

- Endpoint or authentication settings are invalid

In production code:

- Catch `NutrientException`.

- Return a clear error message.

- Log failure details for debugging.

- Add retry logic for transient local server issues.

## Conclusion

Use this workflow to generate image descriptions with local AI:

1. Open the image file with a `using` statement for automatic resource cleanup.

2. The SDK supports multiple image formats, including PNG, JPEG, GIF, BMP, and TIFF.

3. Vision API uses OpenAI-compatible endpoints for local VLM servers by default.

4. Default configuration connects to `http://localhost:1234/v1` with model `qwen/qwen3-vl-4b`.

5. Supported local VLM servers include LM Studio, Ollama, vLLM, and custom inference servers.

6. Optionally configure custom server settings through `Settings.CustomVlmApiSettings`.

7. Create a vision instance with `Vision.Set()` bound to the document for local AI processing.

8. Generate the description with `vision.Describe()`, which sends the image to the local server endpoint and returns natural language text.

9. The SDK encodes image data, constructs OpenAI-compatible multimodal requests, and parses responses automatically.

10. Generated descriptions are concise (1–3 sentences), accessible (WCAG-compliant alt text), accurate (observable details only), and model-dependent.

11. Description quality varies by model size — larger models (7B+) produce more nuanced descriptions than smaller variants (4B).

12. Print or save the description for use in accessibility systems, content management, or cataloging workflows.

13. Handle `NutrientException` failures for vision processing issues, including server unavailable, model not loaded, or timeout errors.

For related image workflows, refer to the [.NET SDK guides](https://www.nutrient.io/guides/dotnet/csharp.md).

Download [this ready-to-use sample package](https://www.nutrient.io/downloads/samples/csharp/describe-image-with-local-ai.zip) to explore local AI image description.
---

## Related pages

- [Nutrient .NET SDK extraction guides](/guides/dotnet/csharp/extraction.md)
- [Applying OCR to a PDF page](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf-page.md)
- [Applying OCR to a PDF document](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf.md)
- [Classifying documents](/guides/dotnet/csharp/extraction/classify-document.md)
- [Generating image descriptions using Claude](/guides/dotnet/csharp/extraction/describe-image-with-claude.md)
- [Generating image descriptions using OpenAI](/guides/dotnet/csharp/extraction/describe-image-with-openai.md)
- [Detecting document language](/guides/dotnet/csharp/extraction/detect-document-language.md)
- [Extracting data from images using ICR](/guides/dotnet/csharp/extraction/extract-data-from-image-icr.md)
- [Extracting data from images using OCR](/guides/dotnet/csharp/extraction/extract-data-from-image-ocr.md)
- [Extracting data from images using vision language models](/guides/dotnet/csharp/extraction/extract-data-from-image-vlm.md)
- [Extracting data from specific pages](/guides/dotnet/csharp/extraction/extract-data-from-specific-pages.md)
- [Extracting form fields from images](/guides/dotnet/csharp/extraction/extract-form-fields-from-image.md)
- [Extracting structured data from documents](/guides/dotnet/csharp/extraction/extract-structured-data.md)
- [Generating extraction schemas](/guides/dotnet/csharp/extraction/generate-extraction-schema.md)
- [Extracting JSON data from a PDF document](/guides/dotnet/csharp/extraction/json-data-extraction.md)
- [Labeling form fields with a vision language model](/guides/dotnet/csharp/extraction/label-form-fields-with-vlm.md)
- [Parsing a document into structured content](/guides/dotnet/csharp/extraction/parse-document.md)
- [Extracting text from PDF documents](/guides/dotnet/csharp/extraction/pdf-to-text.md)
- [Reading barcodes with vision extraction](/guides/dotnet/csharp/extraction/read-barcodes-with-vision.md)
- [Extracting text from multilingual images](/guides/dotnet/csharp/extraction/read-text-from-image-multi-language.md)
- [Extracting text from images](/guides/dotnet/csharp/extraction/read-text-from-image.md)
- [Searching document text](/guides/dotnet/csharp/extraction/search-document-text.md)
- [Speeding up first ICR operation by predownloading models](/guides/dotnet/csharp/extraction/speed-up-first-icr-by-downloading-requirements.md)

