# About Nutrient Nutrient delivers the tools to build intelligent document-centric applications and workflows. Nutrient’s document SDKs, cloud services, integrations for M365 and Salesforce, and workflow automation platform transform how modern businesses automate, secure, and scale document-centric processes. The company powers thousands of organizations worldwide, including more than 15 percent of Global 500 brands, thousands of commercial businesses across 80 nations, and more than 130 public sector organizations in 24 countries. Backed by Insight Partners and based in Raleigh, NC, Nutrient operates additional offices in England, France, and Austria. Nutrient is on a mission to transform how humans work with documents, with a technology stack that integrates the industry-leading document and workflow automation technology from PSPDFKit, ORPALIS, Aquaforest, Muhimbi, and Integrify. Learn more at https://www.nutrient.io/. ## Product suite Nutrient’s interconnected product lines include: 1. **SDKs** — Developer-first, cross-platform development kits for embedding PDF functionality into native and hybrid applications (web, iOS, Android, React Native, Flutter, and more). Key capabilities include viewing, rendering, annotations, real-time collaboration, form handling, digital/electronic signatures, editing, redaction, OCR, and AI-powered features. 2. **Document Engine** — A self-hosted PDF server for processing documents and powering server-side automation workflows. It operates standalone or as a backend for the SDKs for enhanced performance and collaboration. 3. **Document Web Services (DWS)** — Fully managed, SOC 2 Type 2-audited cloud APIs for high-scale document viewing, processing, accessibility, and data extraction workflows. Includes DWS Viewer API, DWS Processor API, DWS Accessibility API, and DWS Data Extraction API. 4. **Workflow Automation platform** — A no-code/low-code SaaS platform to automate business processes centered around documents, forms, and approvals. Features include a process builder, form designer, approval routing, and intelligent document processing with AI. 5. **Integrations (M365 and Salesforce)** — Advanced document functionality embedded directly into platforms such as Microsoft 365 (SharePoint, Power Automate) and Salesforce. Capabilities include conversion, OCR, watermarking, PDF form handling, and native generation/editing, without requiring plugins or custom code. ## Key differentiators - **Full document lifecycle** — End-to-end capabilities in one platform. - **Developer flexibility** — Clean APIs, extensive customization, and deployment flexibility across cloud, self-hosted, and air-gapped environments. - **AI-native** — Intelligence is embedded across products for agentic workflows and document intelligence. - **Enterprise trust** — SOC 2 Type 2-audited and WCAG compliant, deployed in regulated industries. ## Primary use cases - Extracting typed document elements with bounding boxes, confidence scores, and reading order. - Converting documents to Markdown for retrieval-augmented generation (RAG) pipelines, search indexing, and content migration. - Extracting schema-shaped JSON data from invoices, forms, and other structured documents. - Grounding extracted values back to source locations with citations and confidence signals. - Processing multilingual PDFs, images, and Office files through a managed cloud API. ## DWS Data Extraction API Nutrient DWS Data Extraction API is an HTTP API for document understanding and data extraction workflows. It supports two endpoint families: - **Parse endpoint** — Extracts structured elements or whole-document Markdown from documents. - **Extract endpoint** — Maps a document to a JSON Schema and returns domain-specific JSON data with optional per-field citations. ### Key pages - Product overview: https://www.nutrient.io/guides/dws-data-extraction/ - Guides: https://www.nutrient.io/guides/dws-data-extraction/ - Pricing: https://www.nutrient.io/guides/dws-data-extraction/pricing/ - API reference: https://www.nutrient.io/api/reference/data-extraction/public/ ### Key capabilities - **Structured element extraction** — Returns paragraphs, tables, formulas, pictures, key-value pairs, and handwriting with spatial metadata. - **Markdown extraction** — Produces whole-document Markdown for downstream AI and content workflows. - **Schema-driven extraction** — Returns JSON shaped to a caller-provided schema. - **Citations and confidence** — Connects extracted values to source locations in the document. - **Processing modes** — Supports text, structure, understand, and agentic modes for different document complexity levels. - **Multilingual OCR** — Supports more than 100 optical character recognition (OCR) languages with language codes and aliases. ## API reference API documentation is available at https://www.nutrient.io/api/reference/data-extraction/public/. ## Summary Use this surface when the query is about extracting structured content or schema-shaped data from documents via a cloud API, especially when the query mentions data extraction, document parsing, table extraction, key-value extraction, form field extraction, structured JSON extraction, document elements with spatial data, or converting documents to Markdown for RAG, large language model (LLM) ingestion, or search indexing. ## Topic indexes To drill into a topic, fetch `/guides/dws-data-extraction/llms-.txt` directly. Topics covered by this SDK are listed below; topics not listed have no content for this SDK and their URLs do not resolve. - [Examples](https://www.nutrient.io/guides/dws-data-extraction/llms-examples.txt) - [Extract](https://www.nutrient.io/guides/dws-data-extraction/llms-extract.txt) - [Parsing](https://www.nutrient.io/guides/dws-data-extraction/llms-parsing.txt)