Generating image descriptions using OpenAI
Use OpenAI-powered image description to generate alt text and visual summaries in cloud workflows.
Common use cases include:
- Accessibility pipelines for screen readers
- Content management and image cataloging
- Document workflows across regions
- Enterprise integrations with managed API infrastructure
- Fast prototyping without local model hosting
This guide uses OpenAI as the VLM provider through Nutrient Vision API.
Download sampleHow Nutrient helps
Nutrient .NET SDK handles provider configuration, request handling, and response parsing.
The SDK handles:
- OpenAI API authentication and endpoint setup
- Image encoding and multimodal payload formatting
- Model parameters such as temperature and token limits
- API failure and rate-limit handling
Complete implementation
This example generates an image description using OpenAI:
using Nutrient;Configuring the OpenAI provider
Open the image with a using statement(opens in a new tab) and configure OpenAI as the provider.
In this sample:
- Setting
VisionSettings.ProvidertoVlmProvider.OpenAIselects OpenAI. OpenAIApiEndpointSettings.ApiKeysets your API key.- Input can be PNG, JPEG, GIF, BMP, or TIFF.
try{ using Document document = Document.Open("input_photo.png"); var settings = document.Settings;
// Configure OpenAI as the VLM provider settings.VisionSettings.Provider = VlmProvider.OpenAI;
// Set the OpenAI API key settings.OpenAIApiEndpointSettings.ApiKey = "OPENAI_API_KEY";Creating a vision instance and generating the description
Create a vision instance and call Describe() to generate text.
In this sample:
Vision.Set(document)binds processing to the opened image.vision.Describe()returns a description string.- The SDK handles encoding, request construction, and response parsing.
var vision = Vision.Set(document); string description = vision.Describe();Outputting the description
Print the description for review, or store it in your application.
Common destinations include:
- Database fields
- JSON output files
- HTML
altattributes
Console.WriteLine("Image description:"); Console.WriteLine(description);}catch (NutrientException e){ Console.Error.WriteLine($"Error: {e.Message}"); Environment.Exit(1);}Understanding the output
Describe() returns natural language text for accessibility and content understanding.
Descriptions are typically:
- Concise — Focused on key subjects and details, often one to three sentences
- Accessible — Suitable for users who rely on screen readers
- Accurate — Based on visible content only
- Contextual — Include relevant relationships and scene context
Use this output for accessibility metadata, image search, and document workflows.
OpenAI API settings
The OpenAI provider uses these OpenAI API endpoint settings:
ApiEndpoint— The OpenAI API endpoint (default:https://api.openai.com/v1)ApiKey— Your OpenAI API key for authenticationModel— The model identifier to useTemperature— Controls response creativity (0.0 = deterministic, 1.0 = creative)MaxTokens— Maximum tokens in the response (default: 16384)
Error handling
The SDK throws a NutrientException when vision operations fail.
Common failure scenarios include:
- The input image can’t be read due to path, permission, or format issues
- The OpenAI API key is missing or invalid
- The OpenAI API is unavailable
- Rate limits are exceeded
- Network requests fail before reaching the API
- Image data is too large or corrupted
In production code:
- Catch
NutrientException. - Return a clear error message.
- Log failure details for debugging.
- Add retry logic for transient API failures.
Conclusion
Use this workflow to generate image descriptions with OpenAI:
- Open the image file with a
usingstatement for automatic resource cleanup. - The SDK supports multiple image formats, including PNG, JPEG, GIF, BMP, and TIFF.
- Access the vision settings with
Settings.VisionSettingsto configure the VLM provider. - Set
ProvidertoVlmProvider.OpenAIto select OpenAI instead of alternatives like Claude or local models. - Access OpenAI-specific settings with the
OpenAIApiEndpointSettingsproperty for API configuration. - Set the OpenAI API key with the
ApiKeyproperty using credentials obtained from the OpenAI platform. - OpenAI API settings control endpoint URLs, model selection, temperature, and max tokens.
- Create a vision instance with
Vision.Set()bound to the document with configured provider settings. - Generate the description with
vision.Describe(), which sends the image to OpenAI’s vision endpoint and returns natural language text. - The SDK encodes image data, constructs multimodal API requests, and parses responses automatically.
- Generated descriptions are concise (1–3 sentences), accessible (WCAG-compliant alt text), accurate (observable details only), and contextual.
- Print or save the description for use in accessibility systems, content management, or cataloging workflows.
- Handle
NutrientExceptionfailures for vision processing issues, including authentication errors, API failures, and rate limits.
For related image workflows, refer to the .NET SDK guides.
Download this ready-to-use sample package to explore OpenAI-based image description.