This HTML page is not optimized for LLM or AI agent consumption. Fetch the Markdown version instead: /api/csharp/vision.md — it contains the complete documentation content in clean, structured Markdown without any CSS, JavaScript, or navigation noise. Vision

Provides machine learning and computer vision capabilities for document processing. Enables AI-powered document description and content extraction.

using Nutrient;

The SDK creates this class through factory methods or other SDK objects.

Methods

Classify

public string Classify(ClassificationRequest request)

Classifies the document against the candidate labels carried by request (zero-shot) and exports the ranked result. Branch weights, supplied text, and the other knobs are read from DocumentClassificationSettings on the document’s settings. The result is always JSON; OutputFormat is not consulted because only the JSON exporter serializes the classification result.

Parameters

NameTypeDescription
requestClassificationRequestThe classification request carrying at least two candidate labels.

Returns: string - The exported content as a string (JSON), including the predicted label and ranked candidate probabilities.

ClassifyText

public static string ClassifyText(ClassificationRequest request, string text)

Classifies caller-supplied text against the candidate labels carried by request (zero-shot) with no document. The text is scored directly by the text branch and the image branch is skipped, so nothing is rendered and no file is opened — use this when you already have the text (an email body, a database field, your own extraction pipeline) rather than a document on disk.

Parameters

NameTypeDescription
requestClassificationRequestThe classification request carrying at least two candidate labels.
textstringThe text to classify.

Returns: string - JSON with the predicted label and the ranked candidate probabilities.

ClassifyToFile

public void ClassifyToFile(ClassificationRequest request, string outputPath)

Classifies the document against the candidate labels carried by request (zero-shot) and writes the exported result to a file. The result is always JSON; OutputFormat is not consulted because only the JSON exporter serializes the classification result.

Parameters

NameTypeDescription
requestClassificationRequestThe classification request carrying at least two candidate labels.
outputPathstringPath to the output file.

Describe

public string Describe()

Generates an AI-powered description of the document content.

Returns: string - A string containing the document description.

DetectForms

public string DetectForms()

Detects form fields on the document and exports the result. Output format is determined by OutputFormat (JSON or IR Lite). Each detected field carries its type and bounding box. To also assign AI semantic labels (e.g. “First name”), set FormLabelingSettings.EnableAiLabeling on the document’s settings before calling — no separate method is needed.

Returns: string - The exported content as a string (JSON or IR Lite JSON depending on settings).

DetectFormsToFile

public void DetectFormsToFile(string outputPath)

Detects form fields on the document and writes the exported result to a file. Output format is determined by OutputFormat (JSON or IR Lite).

Parameters

NameTypeDescription
outputPathstringPath to the output file.

DetectLanguages

public string DetectLanguages()

Detects the language and text direction of the document, fully offline, and exports the result. Runs the offline cascade (Tesseract OSD for script → script-model OCR → fastText for the specific language) on every page and reports one detection per page. Output format is determined by OutputFormat.

Returns: string - JSON with the predicted language, text direction, and per-page detections.

DetectLanguagesText

public static string DetectLanguagesText(string text)

Detects the language of caller-supplied text, fully offline, with no document. CLD3 scores the text directly — nothing is rendered and no file is opened. Use this when you already have the text (an email body, a database field, your own extraction pipeline) rather than a document on disk.

Parameters

NameTypeDescription
textstringThe text to detect the language of.

Returns: string - JSON with the predicted language and text direction.

public static string DetectLanguagesText(string text, DocumentSettings settings)

Detects the language(s) of caller-supplied text, fully offline, with explicit settings — raise MaxLanguages to report more than one language in mixed text. Without settings the dominant language only is returned.

Parameters

NameTypeDescription
textstringThe text to detect the language(s) of.
settingsDocumentSettingsDocument settings; MaxLanguages bounds how many languages are reported.

Returns: string - JSON with the predicted language(s) and text direction.

DetectLanguagesToFile

public void DetectLanguagesToFile(string outputPath)

Detects the language and text direction of the document, fully offline, and writes the exported result to a file. Output format is determined by OutputFormat.

Parameters

NameTypeDescription
outputPathstringPath to the output file.

ExtractContent

public string ExtractContent()

Extracts structured content from the document using machine vision processing. The pipeline used is determined by the Engine setting and the output format by OutputFormat.

Returns: string - The exported content as a string (JSON, Markdown, or IR Lite JSON depending on settings).

public string ExtractContent(DocumentLayoutJsonExportSettings settings)

Extracts structured content from the document using machine vision processing with custom export settings. The pipeline used is determined by the Engine setting.

Parameters

NameTypeDescription
settingsDocumentLayoutJsonExportSettingsSettings controlling what to include in the JSON output.

Returns: string - A JSON string containing the extracted content structure.

ExtractContentToFile

public void ExtractContentToFile(string outputPath)

Extracts structured content from the document and writes it to a file. The pipeline used is determined by the Engine setting and the output format by OutputFormat.

Parameters

NameTypeDescription
outputPathstringPath to the output file.
public void ExtractContentToFile(string outputPath, DocumentLayoutJsonExportSettings settings)

Extracts structured content from the document and saves it to a JSON file with custom settings.

Parameters

NameTypeDescription
outputPathstringPath to the output JSON file.
settingsDocumentLayoutJsonExportSettingsSettings controlling what to include in the JSON output.

ExtractStructured

public string ExtractStructured(StructuredExtractionRequest request)

Extracts structured data from the document, shaped to the JSON Schema carried by the request’s {"schema": ...} envelope. The document is first read by the extraction pipeline selected by Engine, then an AI model fills the schema from the recognized content. Provider, model, endpoint, and confidence reporting are driven by AiProcessingSettings on the document’s settings.

Parameters

NameTypeDescription
requestStructuredExtractionRequestThe extraction request carrying the schema envelope (required) and optional instructions.

Returns: string - A JSON string with two top-level nodes: extraction (the schema-shaped extracted fields) and metadata (per-field source locations and grounding labels).

ExtractStructuredToFile

public void ExtractStructuredToFile(StructuredExtractionRequest request, string outputPath)

Extracts structured data from the document, shaped to the JSON Schema carried by the request’s {"schema": ...} envelope, and writes the JSON result to a file. See ExtractStructured for the result shape.

Parameters

NameTypeDescription
requestStructuredExtractionRequestThe extraction request carrying the schema envelope (required) and optional instructions.
outputPathstringPath to the output file.

GenerateSchema

public static string GenerateSchema(SchemaGenerationRequest request)

Generates a JSON Schema for a class of documents using a vision model, with default document settings. The schema is designed for downstream LLM structured-output extraction tasks: it is hard-clamped to the knobs in SchemaGenerationSettings and to the target structured-output dialect selected there, so the same call with the same settings always yields a schema the target accepts.

Parameters

NameTypeDescription
requestSchemaGenerationRequestThe schema generation request: the document type the schema must represent (required), an optional natural-language requirement, and up to five example documents used as grounding samples.

Returns: string - A JSON envelope object with a schema property (the generated JSON Schema) and a constraints array of cross-field rules in standard JsonLogic. The array is empty unless SchemaGenerationSettings.IncludeConstraints is enabled.

public static string GenerateSchema(SchemaGenerationRequest request, DocumentSettings settings)

Generates a JSON Schema for a class of documents using a vision model, with explicit document settings. The schema is designed for downstream LLM structured-output extraction tasks: it is hard-clamped to the knobs in SchemaGenerationSettings and to the target structured-output dialect selected there, so the same call with the same settings always yields a schema the target accepts. The vision model connection is resolved from Provider and the matching provider settings class — the same way Describe is configured.

Parameters

NameTypeDescription
requestSchemaGenerationRequestThe schema generation request: the document type the schema must represent (required), an optional natural-language requirement, and up to five example documents used as grounding samples.
settingsDocumentSettingsDocument settings carrying SchemaGenerationSettings (schema shape and target dialect) and the vision model connection.

Returns: string - A JSON envelope object with a schema property (the generated JSON Schema) and a constraints array of cross-field rules in standard JsonLogic. The array is empty unless SchemaGenerationSettings.IncludeConstraints is enabled.

GenerateSchemaToFile

public static void GenerateSchemaToFile(SchemaGenerationRequest request, string outputPath)

Generates a JSON Schema for a class of documents using a vision model and writes it to a file, with default document settings.

Parameters

NameTypeDescription
requestSchemaGenerationRequestThe schema generation request: the document type the schema must represent (required), an optional natural-language requirement, and up to five example documents used as grounding samples.
outputPathstringPath to the output file receiving the JSON Schema.
public static void GenerateSchemaToFile(SchemaGenerationRequest request, DocumentSettings settings, string outputPath)

Generates a JSON Schema for a class of documents using a vision model and writes it to a file, with explicit document settings.

Parameters

NameTypeDescription
requestSchemaGenerationRequestThe schema generation request: the document type the schema must represent (required), an optional natural-language requirement, and up to five example documents used as grounding samples.
settingsDocumentSettingsDocument settings carrying SchemaGenerationSettings (schema shape and target dialect) and the vision model connection.
outputPathstringPath to the output file receiving the JSON Schema.

Set

public static Vision Set(Document document)

Creates a Vision instance for the specified document.

Parameters

NameTypeDescription
documentDocumentThe document to analyze using vision capabilities.

Returns: Vision - A Vision instance ready to perform analysis on the document. Throws: NutrientException - Thrown when document is null.

Split

public string Split()

Splits a merged document into its constituent sub-documents (page-stream segmentation) and exports the result. Every page is scored for whether it starts a new sub-document — fusing a per-page image signal with the page’s text — and the contiguous page ranges are returned. The threshold and the other knobs are read from DocumentSplitSettings on the document’s settings. The result is always JSON; OutputFormat is not consulted because only the JSON exporter serializes the split result.

Returns: string - The exported content as a string (JSON), carrying the detected sub-document segments (1-based inclusive page ranges with their boundary confidence).

SplitToFile

public void SplitToFile(string outputPath)

Splits a merged document into its constituent sub-documents (page-stream segmentation) and writes the exported result to a file. The result is always JSON; OutputFormat is not consulted because only the JSON exporter serializes the split result.

Parameters

NameTypeDescription
outputPathstringPath to the output file.

Warmup

public void Warmup()

Preloads (warms up) all resources needed for vision processing. This downloads all model files based on the document’s VisionSettings before execution. Call this to avoid download delays during ExtractContent().

Resource management

public void Dispose()

Vision implements IDisposable and owns native resources. Always call Dispose() or use a C# using declaration: it is not released automatically.