This HTML page is not optimized for LLM or AI agent consumption. Fetch the Markdown version instead: /api/python/vision.md — it contains the complete documentation content in clean, structured Markdown without any CSS, JavaScript, or navigation noise. Vision

Provides machine learning and computer vision capabilities for document processing. Enables AI-powered document description and content extraction.

from nutrient_sdk import Vision

Construction

Vision cannot be instantiated directly. Obtain instances through static factory methods or via other SDK classes.

Class Methods

classify_text

@classmethod
def classify_text(cls, request: ClassificationRequest, text: str) -> str

Classifies caller-supplied text against the candidate labels carried by request (zero-shot) with no document. The text is scored directly by the text branch and the image branch is skipped, so nothing is rendered and no file is opened — use this when you already have the text (an email body, a database field, your own extraction pipeline) rather than a document on disk.

Parameters:

NameTypeDescription
requestClassificationRequestThe classification request carrying at least two candidate labels.
textstrThe text to classify.

Returns: str - JSON with the predicted label and the ranked candidate probabilities.


detect_languages_text

@classmethod
def detect_languages_text(cls, text: str) -> str
@classmethod
def detect_languages_text(cls, text: str, settings: DocumentSettings) -> str

Detects the language of caller-supplied text, fully offline, with no document. CLD3 scores the text directly — nothing is rendered and no file is opened. Use this when you already have the text (an email body, a database field, your own extraction pipeline) rather than a document on disk.

Parameters:

NameTypeDescription
textstrThe text to detect the language of.
settings (optional)DocumentSettingsDocument settings; MaxLanguages bounds how many languages are reported.

Returns: str - JSON with the predicted language and text direction.


generate_schema

@classmethod
def generate_schema(cls, request: SchemaGenerationRequest) -> str
@classmethod
def generate_schema(cls, request: SchemaGenerationRequest, settings: DocumentSettings) -> str

Generates a JSON Schema for a class of documents using a vision model, with default document settings. The schema is designed for downstream LLM structured-output extraction tasks: it is hard-clamped to the knobs in SchemaGenerationSettings and to the target structured-output dialect selected there, so the same call with the same settings always yields a schema the target accepts.

Parameters:

NameTypeDescription
requestSchemaGenerationRequestThe schema generation request: the document type the schema must represent (required), an optional natural-language requirement, and up to five example documents used as grounding samples.
settings (optional)DocumentSettingsDocument settings carrying SchemaGenerationSettings (schema shape and target dialect) and the vision model connection.

Returns: str - A JSON envelope object with a schema property (the generated JSON Schema) and a constraints array of cross-field rules in standard JsonLogic. The array is empty unless SchemaGenerationSettings.IncludeConstraints is enabled.


generate_schema_to_file

@classmethod
def generate_schema_to_file(cls, request: SchemaGenerationRequest, output_path: str) -> None
@classmethod
def generate_schema_to_file(cls, request: SchemaGenerationRequest, settings: DocumentSettings, output_path: str) -> None

Generates a JSON Schema for a class of documents using a vision model and writes it to a file, with default document settings.

Parameters:

NameTypeDescription
requestSchemaGenerationRequestThe schema generation request: the document type the schema must represent (required), an optional natural-language requirement, and up to five example documents used as grounding samples.
output_pathstrPath to the output file receiving the JSON Schema.
settings (optional)DocumentSettingsDocument settings carrying SchemaGenerationSettings (schema shape and target dialect) and the vision model connection.

set

@classmethod
def set(cls, document: Document) -> Vision

Creates a Vision instance for the specified document.

Parameters:

NameTypeDescription
documentDocumentThe document to analyze using vision capabilities.

Returns: Vision - A Vision instance ready to perform analysis on the document. Raises:


Methods

classify

def classify(self, request: ClassificationRequest) -> str

Classifies the document against the candidate labels carried by request (zero-shot) and exports the ranked result. Branch weights, supplied text, and the other knobs are read from DocumentClassificationSettings on the document’s settings. The result is always JSON; OutputFormat is not consulted because only the JSON exporter serializes the classification result.

Parameters:

NameTypeDescription
requestClassificationRequestThe classification request carrying at least two candidate labels.

Returns: str - The exported content as a string (JSON), including the predicted label and ranked candidate probabilities.


classify_to_file

def classify_to_file(self, request: ClassificationRequest, output_path: str) -> None

Classifies the document against the candidate labels carried by request (zero-shot) and writes the exported result to a file. The result is always JSON; OutputFormat is not consulted because only the JSON exporter serializes the classification result.

Parameters:

NameTypeDescription
requestClassificationRequestThe classification request carrying at least two candidate labels.
output_pathstrPath to the output file.

describe

def describe(self) -> str

Generates an AI-powered description of the document content.

Returns: str - A string containing the document description.


detect_forms

def detect_forms(self) -> str

Detects form fields on the document and exports the result. Output format is determined by OutputFormat (JSON or IR Lite). Each detected field carries its type and bounding box. To also assign AI semantic labels (e.g. “First name”), set FormLabelingSettings.EnableAiLabeling on the document’s settings before calling — no separate method is needed.

Returns: str - The exported content as a string (JSON or IR Lite JSON depending on settings).


detect_forms_to_file

def detect_forms_to_file(self, output_path: str) -> None

Detects form fields on the document and writes the exported result to a file. Output format is determined by OutputFormat (JSON or IR Lite).

Parameters:

NameTypeDescription
output_pathstrPath to the output file.

detect_languages

def detect_languages(self) -> str

Detects the language and text direction of the document, fully offline, and exports the result. Runs the offline cascade (Tesseract OSD for script → script-model OCR → fastText for the specific language) on every page and reports one detection per page. Output format is determined by OutputFormat.

Returns: str - JSON with the predicted language, text direction, and per-page detections.


detect_languages_to_file

def detect_languages_to_file(self, output_path: str) -> None

Detects the language and text direction of the document, fully offline, and writes the exported result to a file. Output format is determined by OutputFormat.

Parameters:

NameTypeDescription
output_pathstrPath to the output file.

extract_content

def extract_content(self) -> str
def extract_content(self, settings: DocumentLayoutJsonExportSettings) -> str

Extracts structured content from the document using machine vision processing. The pipeline used is determined by the Engine setting and the output format by OutputFormat.

Parameters:

NameTypeDescription
settings (optional)DocumentLayoutJsonExportSettingsSettings controlling what to include in the JSON output.

Returns: str - The exported content as a string (JSON, Markdown, or IR Lite JSON depending on settings).


extract_content_to_file

def extract_content_to_file(self, output_path: str) -> None
def extract_content_to_file(self, output_path: str, settings: DocumentLayoutJsonExportSettings) -> None

Extracts structured content from the document and writes it to a file. The pipeline used is determined by the Engine setting and the output format by OutputFormat.

Parameters:

NameTypeDescription
output_pathstrPath to the output file.
settings (optional)DocumentLayoutJsonExportSettingsSettings controlling what to include in the JSON output.

extract_structured

def extract_structured(self, request: StructuredExtractionRequest) -> str

Extracts structured data from the document, shaped to the JSON Schema carried by the request’s {"schema": ...} envelope. The document is first read by the extraction pipeline selected by Engine, then an AI model fills the schema from the recognized content. Provider, model, endpoint, and confidence reporting are driven by AiProcessingSettings on the document’s settings.

Parameters:

NameTypeDescription
requestStructuredExtractionRequestThe extraction request carrying the schema envelope (required) and optional instructions.

Returns: str - A JSON string with two top-level nodes: extraction (the schema-shaped extracted fields) and metadata (per-field source locations and grounding labels).


extract_structured_to_file

def extract_structured_to_file(self, request: StructuredExtractionRequest, output_path: str) -> None

Extracts structured data from the document, shaped to the JSON Schema carried by the request’s {"schema": ...} envelope, and writes the JSON result to a file. See ExtractStructured for the result shape.

Parameters:

NameTypeDescription
requestStructuredExtractionRequestThe extraction request carrying the schema envelope (required) and optional instructions.
output_pathstrPath to the output file.

split

def split(self) -> str

Splits a merged document into its constituent sub-documents (page-stream segmentation) and exports the result. Every page is scored for whether it starts a new sub-document — fusing a per-page image signal with the page’s text — and the contiguous page ranges are returned. The threshold and the other knobs are read from DocumentSplitSettings on the document’s settings. The result is always JSON; OutputFormat is not consulted because only the JSON exporter serializes the split result.

Returns: str - The exported content as a string (JSON), carrying the detected sub-document segments (1-based inclusive page ranges with their boundary confidence).


split_to_file

def split_to_file(self, output_path: str) -> None

Splits a merged document into its constituent sub-documents (page-stream segmentation) and writes the exported result to a file. The result is always JSON; OutputFormat is not consulted because only the JSON exporter serializes the split result.

Parameters:

NameTypeDescription
output_pathstrPath to the output file.

warmup

def warmup(self) -> None

Preloads (warms up) all resources needed for vision processing. This downloads all model files based on the document’s VisionSettings before execution. Call this to avoid download delays during ExtractContent().