AI Search Pulse
Document AI / IDP Monthly Benchmark: August 2026
Benchmark Date: August 10, 2026
Methodology Version: 1.0
AI Assistant: ChatGPT
Model: GPT-5.6 Luna
Tracked Vendors: 10
Prompt Library: 49 Buyer Evaluation Scenarios
Executive Summary
This is the second benchmark published by AI Search Pulse for the Intelligent Document Processing (IDP) category, following the July 2026 baseline.
Using the methodology described in the AI Search Pulse Research Methodology (Version 1.0), we submitted the same standardized library of 49 buyer-representative prompts used in July to ChatGPT and recorded which enterprise software vendors were recommended, how frequently, and where they appeared within each response.
Important: Between the July and August benchmark runs, OpenAI released GPT-5.6, and standard ChatGPT access shifted to this new model family — specifically GPT-5.6 Luna, OpenAI’s fastest and lowest-cost tier. As a result, this report reflects a different underlying model than the July baseline. Month-over-month figures should be read as directional context, not a controlled comparison. This run was conducted using Temporary Chat with memory and chat history reference disabled, to rule out account-level personalization as a factor.
This benchmark measures AI recommendation visibility, not product quality, customer satisfaction, analyst rankings, or market share.
Key Findings
Before the detailed metrics, four patterns stood out in this benchmark.
1. The enterprise leadership group narrowed from four vendors to three. ABBYY, UiPath, and Hyperscience formed a closer, more consistent group in enterprise platform-selection prompts than in July, when a fourth vendor, Tungsten Automation, sat within the same tight band.
2. Tungsten Automation’s use-case visibility dropped sharply. Use-case and industry-specific mentions for Tungsten Automation fell from a consistent double-digit presence in July to 3 in this run, the largest single-category decline of any tracked vendor.
3. Rossum’s visibility became more evenly distributed across buyer personas. Where July showed a pronounced gap between Rossum’s developer/API visibility and its enterprise visibility, this benchmark shows a more even spread across categories.
4. Some prompts returned no tracked vendors, or answers unrelated to the prompt’s intent. A small number of developer-focused prompts returned generic platform mentions or, in one case, identity/authentication tools unrelated to document processing — a pattern not observed in the July baseline.
AI Visibility Summary
| Vendor | Total Mentions | Enterprise | Use Case / Industry | API | Competitive | AI Visibility Score |
|---|---|---|---|---|---|---|
| ABBYY | 42 | 11 | 10 | 10 | 11 | 18.5% |
| UiPath | 41 | 11 | 10 | 9 | 11 | 18.1% |
| Hyperscience | 38 | 9 | 9 | 9 | 11 | 16.7% |
| Rossum | 31 | 6 | 7 | 9 | 9 | 13.7% |
| Tungsten Automation | 24 | 7 | 3 | 6 | 8 | 10.6% |
| Nanonets | 18 | 5 | 0 | 8 | 5 | 7.9% |
| Docsumo | 9 | 2 | 1 | 3 | 3 | 4.0% |
| Klippa | 9 | 1 | 0 | 5 | 3 | 4.0% |
| Veryfi | 8 | 1 | 0 | 4 | 3 | 3.5% |
| Affinda | 4 | 1 | 0 | 2 | 1 | 1.8% |
AI Visibility Score represents each vendor’s share of total tracked recommendations within the benchmark prompt library.
Benchmark Observations
The top three tightened; Tungsten Automation fell out of the leading group. ABBYY, UiPath, and Hyperscience remained closely grouped, while Tungsten Automation’s score dropped to 10.6%, driven primarily by its use-case/industry decline. Whether this reflects a genuine shift or model-specific behavior cannot be determined from a single data point.
Rossum’s category imbalance narrowed. Rossum’s enterprise mentions rose from 3 (July) to 6 (August) against a similar total mention count, suggesting a more even presence across buyer personas than the baseline showed.
Nanonets’ enterprise gains did not hold in this run. Nanonets returned to zero use-case mentions in this benchmark. Continued tracking will clarify whether this is a reversal or normal variation.
Model behavior differences were directly observable. Several responses in this run returned no tracked vendors, or, in one case, named identity/authentication platforms in response to a document-processing SDK question. This kind of miss was not present in July and is consistent with GPT-5.6 Luna’s positioning as OpenAI’s lowest-capability tier.
Questions for the Next Benchmark
- Does GPT-5.6 Luna remain the standard ChatGPT experience, or does access shift again?
- Does Tungsten Automation’s use-case decline persist under a second Luna-based run?
- Does Rossum’s more balanced category distribution hold?
- Can any changes be more confidently attributed to vendor-specific visibility versus model-specific behavior once two same-model data points exist?
Methodology & Limitations
This benchmark measures AI recommendation visibility within a standardized set of buyer evaluation scenarios. It does not evaluate product quality, analyst rankings, customer satisfaction, or market share.
This report was conducted August 10, 2026, using ChatGPT (GPT-5.6 Luna) via Temporary Chat with memory and chat history reference disabled. It reflects a different underlying model from the July 2026 baseline (GPT-5.5) due to a platform-side model transition outside this project’s control. Month-over-month percentage comparisons should not be read as a controlled measure of vendor-specific visibility change. Any methodological changes are documented through public version history.
AI Search Pulse Roadmap
| Status | Milestone |
|---|---|
| ✅ Published | Research Methodology v1.0 |
| ✅ Published | Document AI / IDP Baseline Benchmark (July 2026) |
| ✅ Published | Document AI / IDP Monthly Benchmark (August 2026) |
| 📅 Planned | September 2026 Monthly Benchmark |
| 📅 Planned | ChatGPT + Gemini Comparative Benchmark |
| 📅 Planned | Additional Enterprise Software Categories |
Want to know how your company shows up in this data? Book your strategy call today



