---
title: "Label form fields with a VLM in C# | Nutrient .NET SDK"
canonical_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/label-form-fields-with-vlm/"
md_url: "https://www.nutrient.io/guides/dotnet/csharp/extraction/label-form-fields-with-vlm.md"
last_updated: "2026-10-08T00:00:00.000Z"
description: "Detect form fields and assign semantic labels using a vision language model with Nutrient .NET SDK."
---

# Labeling form fields with a vision language model

Form-field detection locates the fillable regions on a page, but a bounding box alone doesn't tell you what each field means. AI labeling adds a human-readable semantic label to each detected field, such as "First name" or "Date of birth", by sending the page to a vision language model (VLM).

This sample builds on offline form-field detection. For the detection basics — and a fully offline workflow with no model contacted — refer to the [extract form fields from an image](https://www.nutrient.io/guides/dotnet/csharp/extraction/extract-form-fields-from-image.md) guide. Here you connect a VLM provider and turn labeling on with Nutrient.NET SDK.

[Download sample](https://www.nutrient.io/downloads/samples/csharp/label-form-fields-with-vlm.zip)

## How Nutrient helps

Nutrient.NET SDK runs detection and labeling behind a single method call. With labeling enabled, it also:

- Draws numbered marks over each detected field on a rendered copy of the page

- Sends the annotated page to the VLM you select and reads each field's semantic label back

- Optionally drops detections the model judges to be false positives

- Records each field's type, bounding box, confidence, and assigned label in JSON

The result is structured data you can index, validate, or feed into a downstream workflow.

## Prerequisites

AI labeling requires a reachable VLM endpoint. The SDK does not provision or start a VLM service for you.

- Configure a reachable VLM endpoint in your environment.

- Configure the API endpoint and model in [custom VLM API settings](https://www.nutrient.io/api/csharp/settings/vision/advanced/custom-vlm-api-settings.md).

- By default, the SDK may assume:
  - API endpoint: `http://localhost:1234/v1`
  - Model: `qwen/qwen3-vl-8b`

- For clarity and reliability, set both the API endpoint and the model explicitly.

- Example with [LM Studio](https://lmstudio.ai/):
  - Run LM Studio in server mode.
  - Load a compatible vision model such as Qwen3-VL (4B, 8B, or larger depending on your hardware).

- Make sure the endpoint is running before you call `DetectForms()` with labeling enabled.

If no VLM endpoint is available, labeling fails at runtime. Leave AI labeling at its default of `false` to run detection only and keep the workflow offline.

## Connect a vision model

Labeling uses the same provider configuration as the rest of the Vision API, so you don't configure a separate endpoint for form labeling. Set the provider in [vision settings](https://www.nutrient.io/api/csharp/settings/vision/vision-settings.md#provider) and fill in the matching provider settings class:

- **Custom / local (default)** — An OpenAI-compatible server such as [LM Studio](https://lmstudio.ai/), Ollama, or vLLM. Configure [custom VLM API settings](https://www.nutrient.io/api/csharp/settings/vision/advanced/custom-vlm-api-settings.md).

- **OpenAI** — Configure [OpenAI API endpoint settings](https://www.nutrient.io/api/csharp/settings/vision/advanced/open-ai-api-endpoint-settings.md).

## Complete implementation

Start by importing the Nutrient namespace:

```csharp

using Nutrient;

```

## Load the document

Open the document with a [using statement](https://learn.microsoft.com/dotnet/csharp/language-reference/statements/using) so the handle is released when the block ends:

```csharp

try
{
    using Document document = Document.Open("input_forms_detection.pdf");

```

## Configure AI labeling

Select the provider, point it at your vision model, then opt in by setting `EnableAiLabeling` to `true`:

```csharp

    var settings = document.Settings;

    // Select the vision model provider (the same setting Vision.Describe() uses)
    settings.VisionSettings.Provider = VlmProvider.Custom;

    // Configure the matching provider settings class
    var vlm = settings.CustomVlmApiSettings;
    vlm.ApiEndpoint = "http://localhost:1234/v1";
    vlm.Model = "qwen/qwen3-vl-8b";

    // Turn on labeling and drop detections the model judges to be false positives
    var formLabeling = settings.FormLabelingSettings;
    formLabeling.EnableAiLabeling = true;
    formLabeling.EnableAiRemoveFalsePositives = true;

    // Optional: constrain labels to a known vocabulary
    formLabeling.CandidateLabels = "First name, Last name, Date of birth, Signature";

```

## Detect and label form fields

Create a vision instance from the document with `Vision.Set(document)`, then call `DetectForms()`. The same call covers both modes; it includes labels because AI labeling is enabled:

```csharp

    var vision = Vision.Set(document);
    string formsJson = vision.DetectForms();

```

Write the JSON result to a file for downstream processing:

```csharp

    File.WriteAllText("output.json", formsJson);
}
catch (NutrientException e)
{
    Console.Error.WriteLine($"Error: {e.Message}");
    Environment.Exit(1);
}

```

## Match labels to a vocabulary

Free-form labels can vary between runs ("First name" vs. "Given name"), which makes them hard to map to a database or template. Supply a vocabulary of preferred labels with the `CandidateLabels` property, as shown above, and the model maps each field to one when it fits. If no label fits, it invents a concise new label.

Pass the labels as newline- or comma-separated text. A matched label uses the casing you supplied, and each field's `labelSource` records whether the label was `matched` or `invented`. Leave the candidate labels empty, which is the default, for free-form labeling.

## Understand the output

`DetectForms()` returns structured JSON. The `elements` array holds one form element per page. Each form element includes its `pageNumber` and a `fields` list, so fields from a multi-page document stay grouped by the page they came from. Each field includes:

- **`fieldType`** — The detected type: `Text`, `Checkbox`, or `Signature`.

- **`bounds`** — The bounding box of the field on the page.

- **`confidence`** — The detection confidence for the field.

- **`label`** — The AI-assigned semantic label (for example, "First name"). Present only when AI labeling is enabled.

- **`labelSource`** — `matched` or `invented`, present only when a candidate vocabulary was supplied.

- **`id`** — A unique identifier for the field.

## Handle errors

Vision API throws a `NutrientException` when detection or labeling fails.

Common failure scenarios include:

- The document can't be read due to path or permission issues

- The page produces no renderable image

- The form detection model is missing or inaccessible, or the feature isn't licensed

- AI labeling is enabled but the selected provider's endpoint is unreachable

In production code:

- Catch `NutrientException`.

- Return a clear error message.

- Log failure details for debugging.

- Consider running detection only, with labeling disabled, as a fallback when the vision endpoint is unavailable.

## Conclusion

The workflow for labeling form fields with a vision model is:

1. Open the source document with a `using` statement for automatic resource cleanup.

2. Select a provider with `VisionSettings.Provider`, configure the matching provider settings class, then set `EnableAiLabeling` to `true` on the form labeling settings.

3. Create a vision instance with `Vision.Set()`.

4. Call `DetectForms()` to detect every field, assign a semantic label, and export the result as JSON.

5. Write the JSON to a file for indexing, validation, or downstream processing.

6. Handle `NutrientException` failures for robust error recovery.

Labeling adds semantic meaning when a vision model is available. For offline detection with no model contacted, refer to the [extract form fields from an image](https://www.nutrient.io/guides/dotnet/csharp/extraction/extract-form-fields-from-image.md) guide. To produce a fillable PDF instead of data, refer to the [detect and add form fields](https://www.nutrient.io/guides/dotnet/csharp/editor/detect-and-add-form-fields.md) guide.

For related image extraction workflows, refer to the [.NET SDK](https://www.nutrient.io/guides/dotnet/csharp.md) guides.

Download the [sample package](https://www.nutrient.io/downloads/samples/csharp/label-form-fields-with-vlm.zip) to explore form-field labeling.
---

## Related pages

- [Nutrient .NET SDK extraction guides](/guides/dotnet/csharp/extraction.md)
- [Applying OCR to a PDF page](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf-page.md)
- [Applying OCR to a PDF document](/guides/dotnet/csharp/extraction/apply-ocr-to-pdf.md)
- [Classifying documents](/guides/dotnet/csharp/extraction/classify-document.md)
- [Generating image descriptions using Claude](/guides/dotnet/csharp/extraction/describe-image-with-claude.md)
- [Generating image descriptions using local AI](/guides/dotnet/csharp/extraction/describe-image-with-local-ai.md)
- [Generating image descriptions using OpenAI](/guides/dotnet/csharp/extraction/describe-image-with-openai.md)
- [Detecting document language](/guides/dotnet/csharp/extraction/detect-document-language.md)
- [Extracting data from images using ICR](/guides/dotnet/csharp/extraction/extract-data-from-image-icr.md)
- [Extracting data from images using OCR](/guides/dotnet/csharp/extraction/extract-data-from-image-ocr.md)
- [Extracting data from images using vision language models](/guides/dotnet/csharp/extraction/extract-data-from-image-vlm.md)
- [Extracting data from specific pages](/guides/dotnet/csharp/extraction/extract-data-from-specific-pages.md)
- [Extracting form fields from images](/guides/dotnet/csharp/extraction/extract-form-fields-from-image.md)
- [Extracting structured data from documents](/guides/dotnet/csharp/extraction/extract-structured-data.md)
- [Generating extraction schemas](/guides/dotnet/csharp/extraction/generate-extraction-schema.md)
- [Extracting JSON data from a PDF document](/guides/dotnet/csharp/extraction/json-data-extraction.md)
- [Parsing a document into structured content](/guides/dotnet/csharp/extraction/parse-document.md)
- [Extracting text from PDF documents](/guides/dotnet/csharp/extraction/pdf-to-text.md)
- [Reading barcodes with vision extraction](/guides/dotnet/csharp/extraction/read-barcodes-with-vision.md)
- [Extracting text from multilingual images](/guides/dotnet/csharp/extraction/read-text-from-image-multi-language.md)
- [Extracting text from images](/guides/dotnet/csharp/extraction/read-text-from-image.md)
- [Searching document text](/guides/dotnet/csharp/extraction/search-document-text.md)
- [Speeding up first ICR operation by predownloading models](/guides/dotnet/csharp/extraction/speed-up-first-icr-by-downloading-requirements.md)

