AI Search Pulse: Document AI / IDP Report: August 2026

 

AI Search Pulse

Document AI / IDP Monthly Benchmark: August 2026

Benchmark Date: August 10, 2026

Methodology Version: 1.0

AI Assistant: ChatGPT

Model: GPT-5.6 Luna

Tracked Vendors: 10

Prompt Library: 49 Buyer Evaluation Scenarios

Executive Summary

This is the second benchmark published by AI Search Pulse for the Intelligent Document Processing (IDP) category, following the July 2026 baseline.

Using the methodology described in the AI Search Pulse Research Methodology (Version 1.0), we submitted the same standardized library of 49 buyer-representative prompts used in July to ChatGPT and recorded which enterprise software vendors were recommended, how frequently, and where they appeared within each response.

Important: Between the July and August benchmark runs, OpenAI released GPT-5.6, and standard ChatGPT access shifted to this new model family — specifically GPT-5.6 Luna, OpenAI’s fastest and lowest-cost tier. As a result, this report reflects a different underlying model than the July baseline. Month-over-month figures should be read as directional context, not a controlled comparison. This run was conducted using Temporary Chat with memory and chat history reference disabled, to rule out account-level personalization as a factor.

This benchmark measures AI recommendation visibility, not product quality, customer satisfaction, analyst rankings, or market share.

Key Findings

Before the detailed metrics, four patterns stood out in this benchmark.

1. The enterprise leadership group narrowed from four vendors to three. ABBYY, UiPath, and Hyperscience formed a closer, more consistent group in enterprise platform-selection prompts than in July, when a fourth vendor, Tungsten Automation, sat within the same tight band.

2. Tungsten Automation’s use-case visibility dropped sharply. Use-case and industry-specific mentions for Tungsten Automation fell from a consistent double-digit presence in July to 3 in this run, the largest single-category decline of any tracked vendor.

3. Rossum’s visibility became more evenly distributed across buyer personas. Where July showed a pronounced gap between Rossum’s developer/API visibility and its enterprise visibility, this benchmark shows a more even spread across categories.

4. Some prompts returned no tracked vendors, or answers unrelated to the prompt’s intent. A small number of developer-focused prompts returned generic platform mentions or, in one case, identity/authentication tools unrelated to document processing — a pattern not observed in the July baseline.

AI Visibility Summary

Vendor Total Mentions Enterprise Use Case / Industry API Competitive AI Visibility Score
ABBYY 42 11 10 10 11 18.5%
UiPath 41 11 10 9 11 18.1%
Hyperscience 38 9 9 9 11 16.7%
Rossum 31 6 7 9 9 13.7%
Tungsten Automation 24 7 3 6 8 10.6%
Nanonets 18 5 0 8 5 7.9%
Docsumo 9 2 1 3 3 4.0%
Klippa 9 1 0 5 3 4.0%
Veryfi 8 1 0 4 3 3.5%
Affinda 4 1 0 2 1 1.8%

AI Visibility Score represents each vendor’s share of total tracked recommendations within the benchmark prompt library.

Benchmark Observations

The top three tightened; Tungsten Automation fell out of the leading group. ABBYY, UiPath, and Hyperscience remained closely grouped, while Tungsten Automation’s score dropped to 10.6%, driven primarily by its use-case/industry decline. Whether this reflects a genuine shift or model-specific behavior cannot be determined from a single data point.

Rossum’s category imbalance narrowed. Rossum’s enterprise mentions rose from 3 (July) to 6 (August) against a similar total mention count, suggesting a more even presence across buyer personas than the baseline showed.

Nanonets’ enterprise gains did not hold in this run. Nanonets returned to zero use-case mentions in this benchmark. Continued tracking will clarify whether this is a reversal or normal variation.

Model behavior differences were directly observable. Several responses in this run returned no tracked vendors, or, in one case, named identity/authentication platforms in response to a document-processing SDK question. This kind of miss was not present in July and is consistent with GPT-5.6 Luna’s positioning as OpenAI’s lowest-capability tier.

Questions for the Next Benchmark

  • Does GPT-5.6 Luna remain the standard ChatGPT experience, or does access shift again?
  • Does Tungsten Automation’s use-case decline persist under a second Luna-based run?
  • Does Rossum’s more balanced category distribution hold?
  • Can any changes be more confidently attributed to vendor-specific visibility versus model-specific behavior once two same-model data points exist?

Methodology & Limitations

This benchmark measures AI recommendation visibility within a standardized set of buyer evaluation scenarios. It does not evaluate product quality, analyst rankings, customer satisfaction, or market share.

This report was conducted August 10, 2026, using ChatGPT (GPT-5.6 Luna) via Temporary Chat with memory and chat history reference disabled. It reflects a different underlying model from the July 2026 baseline (GPT-5.5) due to a platform-side model transition outside this project’s control. Month-over-month percentage comparisons should not be read as a controlled measure of vendor-specific visibility change. Any methodological changes are documented through public version history.

AI Search Pulse Roadmap

Status Milestone
✅ Published Research Methodology v1.0
✅ Published Document AI / IDP Baseline Benchmark (July 2026)
✅ Published Document AI / IDP Monthly Benchmark (August 2026)
📅 Planned September 2026 Monthly Benchmark
📅 Planned ChatGPT + Gemini Comparative Benchmark
📅 Planned Additional Enterprise Software Categories

Want to know how your company shows up in this data? Book your strategy call today

 

Tags
Share on:

Table of Contents

Related Posts