Changelog for the Data Extraction API
RSS2026.10.1 - 1 Oct 2026
- ChangedImproves table extraction for dropped or split tables, wrapped rows, stacked records, shaded regions, merged cells, and false table detection.
- ChangedImproves scanned-document accuracy by running optical character recognition (OCR) alongside vision model analysis and resolving disagreements with evidence fusion. (L#NAVI-124, L#NAVI-125, L#NAVI-130)
- ChangedImproves OCR language handling so Latin scans are less likely to be misclassified as Cyrillic, Arabic, or Greek, and more configured OCR languages are honored. (L#NAVI-148, L#NAPY-24)
- ChangedCharts now return their data as tables by default, and QR codes and barcodes now return their decoded content.
2026.9.21 - 21 Sep 2026
- ChangedIncreases extract schema limits: schema size from 32 KB to 128 KB, total fields from 500 to 2,000, properties per object from 50 to 200, required entries from 50 to 200, and description length from 1,024 to 4,096 characters. (#58687)
2026.9.3 - 3 Sep 2026
- ChangedClarifies the classify API documentation and specification: Scores are independent, and non-zero text/image weights are ignored by the engine. (#58039)
2026.8.21 - 21 Aug 2026
- ChangedUpdates parse to return an actionable 400 error for password-protected documents instead of a generic 500. (#57421)
2026.8.7 - 7 Aug 2026
- FixedFixes classify reporting placeholder
pagesProcessedvalues, and fixes billing per request instead of per processed page. (#56853)
2026.6.5 - 5 Jun 2026
- ChangedUpdates extract billing so requests are billed as parse and extract components. Responses now include a
price_compositionbreakdown. (#53930)