Settings for Extract. Values fall back through three levels: document → SDK → built-in default. Writes target the document only when set on a document’s settings, otherwise the SDK globally when set on SdkSettings.
using Nutrient;The SDK creates this class through factory methods or other SDK objects.
Properties
ApiKey
public string? ApiKey { get; set; }API key for the configured provider. Legacy flat connection (see Provider).
Type: string?
Endpoint
public string? Endpoint { get; set; }Endpoint URL. Required for the local provider. Optional for openai — supply it to route through an OpenAI-compatible proxy/gateway; the public OpenAI endpoint is used when unset. Optional for anthropic (defaults to the public Anthropic base when unset). When set, must be a valid absolute URL. Legacy flat connection (see Provider).
Type: string?
ExtractionSchema
public string? ExtractionSchema { get; set; }Gets or sets the value of ExtractionSchema. Reads walk the chain: Document override → SDK override → Default. Writes go to document level if available, otherwise SDK level.
Type: string?
ExtractionSchemaFile
public string? ExtractionSchemaFile { get; set; }Gets or sets the value of ExtractionSchemaFile. Reads walk the chain: Document override → SDK override → Default. Writes go to document level if available, otherwise SDK level.
Type: string?
IncludeCompositeConfidence
public bool IncludeCompositeConfidence { get; set; }(Schema v1 only.) Daemon-level default for whether the legacy rolled-up composite number is attached to result.json metadata leaves as confidence. No effect under schema v2 (the composite is internal-only there). Defaults to false; a per-request includeCompositeConfidence action parameter overrides it.
Type: bool
IncludeConfidence
public bool IncludeConfidence { get; set; }Daemon-level default for whether AI processing capacities emit confidence output - the confidence.json sidecar and the per-field confidence signals on result.json metadata. Defaults to false (off); a per-request includeConfidence action parameter overrides it.
Type: bool
IncludeDiagnostics
public bool? IncludeDiagnostics { get; set; }Gets or sets the value of IncludeDiagnostics. Reads walk the chain: Document override → SDK override → Default. Writes go to document level if available, otherwise SDK level.
Type: bool?
IncludePageImages
public bool IncludePageImages { get; set; }Opt-in (default false): when true, every page of the source document is rendered to an image and sent to the model alongside the IR-lite text (multimodal extraction). The text stays authoritative for source grounding; images add a visual signal. There is no page cap — large documents are fit to the model’s context window by windowing across multiple calls, not by dropping pages. A per-request includePageImages action parameter overrides it.
Type: bool
IncludeSourceLocations
public bool IncludeSourceLocations { get; set; }Whether extraction grounds each extracted field back to its source location in the document — the metadata node of result.json with per-field match labels and source bounding boxes. Defaults to true. Turn off to cut model token usage when only the extracted values matter: grounding mirrors the schema in the structured-output request so the model also returns per-field source ids. A per-request includeSourceLocations action parameter overrides it.
Type: bool
Instructions
public string? Instructions { get; set; }Gets or sets the value of Instructions. Reads walk the chain: Document override → SDK override → Default. Writes go to document level if available, otherwise SDK level.
Type: string?
MaxAttempts
public int? MaxAttempts { get; set; }Maximum retry attempts on structured-output failure. null uses the capacity default of 1. Provider-independent — applies regardless of how the connection is configured.
Type: int?
MaxInputTokens
public int MaxInputTokens { get; set; }Input-token budget per model call. When the estimated input (schema + IR-lite text + page images) exceeds this, the document is split into page-level windows that each fit, extracted independently, and merged — so a large document never hard-fails on the provider’s context limit. Defaults to 100000, a conservative value that fits common model context windows while leaving room for output; raise it for large-context models. A per-request maxInputTokens action parameter overrides it.
Type: int
MaxParallelCalls
public int MaxParallelCalls { get; set; }Maximum number of windowed extraction calls to run concurrently. Only applies when a large document is split into page-level windows (see MaxInputTokens); each window is an independent VLM call, so running several at once is a large wall-clock win. Defaults to 4 — a balance between throughput and provider rate limits; raise it for large-context keys, lower it (or set 1 for fully sequential) if you hit rate limits. A per-request maxParallelCalls action parameter overrides it.
Type: int
Model
public string? Model { get; set; }Model identifier (e.g. "gpt-4o" or a local model id). Legacy flat connection (see Provider).
Type: string?
OutputDirectory
public string? OutputDirectory { get; set; }Gets or sets the value of OutputDirectory. Reads walk the chain: Document override → SDK override → Default. Writes go to document level if available, otherwise SDK level.
Type: string?
OutputFilePath
public string? OutputFilePath { get; set; }Gets or sets the value of OutputFilePath. Reads walk the chain: Document override → SDK override → Default. Writes go to document level if available, otherwise SDK level.
Type: string?
Provider
public string? Provider { get; set; }Provider discriminator. One of "openai", "local" (custom OpenAI-compatible endpoint), or "anthropic" (alias "claude", native Anthropic Messages API).
Type: string?
ReasoningEffort
public string? ReasoningEffort { get; set; }Gets or sets the value of ReasoningEffort. Reads walk the chain: Document override → SDK override → Default. Writes go to document level if available, otherwise SDK level.
Type: string?
StrictStructuredOutput
public bool StrictStructuredOutput { get; set; }Default true: structured output runs in the provider’s strict mode — the response is grammar-constrained to the schema, which is normalized automatically (additionalProperties: false, all properties required, unsupported keywords moved into descriptions). On by default because the schema is otherwise only advisory: without strict, the provider compiles no decoding grammar and nothing stops the model returning a payload that doesn’t match, which is what produced “AI processing failed to produce valid JSON” failures in production. Note that under strict mode the model must emit every schema property — absence stays expressible because optional properties are made nullable before the schema reaches the provider. Set to false, or pass the per-request strictStructuredOutput action parameter, to opt back out.
Type: bool
Temperature
public double? Temperature { get; set; }Sampling temperature passed to the model. null uses the provider default. Legacy flat connection (see Provider).
Type: double?
Resource management
public void Dispose()ExtractSettings implements IDisposable. Call Dispose() or use a C# using declaration to release its native handle right away. If you don’t, the handle is released when the object is garbage collected.