9  Vendor Selection

NoteExecutive summary
  • Vendor selection is a stack of four orthogonal questions — volume, document variability, compliance posture, self-host vs. SaaS — that collapse the vendor space by roughly 80% before benchmark scores enter the conversation. Start there, not with a vendor’s marketing site.
  • Healthcare prior-authorization (post-CMS rule), legal contract review, KYC document review, and financial filings are the clearest 2026 ROI cases for Pattern C agentic OCR. Generic invoice processing, commodity document indexing, and any sub-cent-per-page workload typically are not. Section 9.5 details the “should I build this at all” question with explicit cost math.
  • Buy the parser, build the eval. The single most important principle in vendor selection. Vendors can do parsing well; only you can define correctness for your documents.
  • A shortlist of three vendors across three tiers, evaluated on a 200–500 document golden set drawn from your production traffic, is the methodology that consistently produces good decisions. Anything cheaper than that is gambling.

9.1 A decision framework

flowchart TD
    A[Start: define use case] --> B{Volume?}
    B -- Low (<10k pages/mo) --> C[Hyperscaler API]
    B -- Medium --> D{Doc variety?}
    B -- High --> E{Self-host?}
    D -- Narrow --> F[Classic IDP or template-based]
    D -- Wide --> G[AI-native specialist]
    E -- Yes --> H[Open-source agentic stack]
    E -- No --> G

Vendor-selection decision tree

The decision tree above is the compressed version of this chapter, and you should be able to use it on its own without reading further. The expansion that follows is for the cases where the compressed answer is not quite right for your situation — which is most cases, because vendor selection is genuinely workload-specific.

The four orthogonal questions, in the order they should be asked:

Volume. Below 10,000 pages per month, a hyperscaler API in Pattern A mode is almost always the right answer regardless of the rest of the analysis. The math does not support investing in Pattern B or C — and certainly not in a vendor relationship — until volume justifies it. Between 10,000 and 1,000,000 pages per month, the AI-native specialists and Pattern B architectures dominate; this is the 2026 sweet spot for the “buy from a vendor” answer. Above 1,000,000 pages per month, self-hosted open-source becomes economically compelling, and engineering capacity becomes the binding constraint rather than vendor pricing.

Document variability. Narrow and stable (one vendor template, three field types, layouts changing once per year) → template-driven classic IDP can still win on TCO. Wide and unstable (mixed vendors, evolving formats, occasional novel document types) → AI-native specialists with Pattern B or C architectures. The mistake is paying for variability handling you do not have; the mirror mistake is choosing template-driven IDP when your variability is wider than the templates can cover, and discovering this at scale.

Compliance posture. HIPAA, SOC 2 Type II, GDPR DPA, FedRAMP, ZDR (Zero Data Retention), data residency requirements. These collapse the vendor space sharply. If you need HIPAA with a BAA, you immediately rule out vendors that do not offer one. If you need on-premises deployment for data sovereignty reasons, you rule out SaaS-only vendors and most of the hyperscaler offerings. Compliance is rarely the most discussed criterion in technical decks; it is often the most consequential in actual procurement.

Two additional 2026 compliance filters worth naming explicitly because they have become procurement-decisive faster than the older certifications:

  • EU AI Act. In force since August 2024, with phased compliance obligations rolling through 2027. Document AI systems used in the EU for “high-risk” purposes (employment decisions, access to essential services, law enforcement, migration, justice administration) face the heaviest obligations — risk management, data governance, transparency, human oversight, and post-market monitoring. For most enterprise document workflows the obligations are lighter, but the act’s transparency requirements still apply (output must be identifiable as AI-generated where users would otherwise reasonably assume it is human-generated). Vendor selection in 2026 increasingly involves an “AI Act conformity” question alongside the older GDPR DPA question; vendors with substantial EU customer bases (Mistral, ABBYY, Klippa, Rossum) are typically ahead on this.
  • FedRAMP authorization. Required for US federal government workloads. The three hyperscalers (AWS GovCloud, Azure Government, GCP Assured Workloads) all hold FedRAMP High. The AI-native specialists are mostly not FedRAMP-authorized in 2026 — a meaningful gap for buyers in federal civilian, defense, or intelligence space. Some specialists offer GovCloud-deployed instances via the hyperscaler relationship, which can be a workable bridge.

Self-host vs. SaaS. Closely related to compliance but distinct. Some teams have data-sovereignty requirements that mandate self-hosting; others have engineering capacity that makes self-hosting attractive on economic grounds; others have neither and should not self-host. The honest framing: self-hosting open-source agentic OCR is not “cheaper” by default. It is cheaper if your team can operate it, and most teams discover the operational cost is higher than they expected.

These four questions, answered before any vendor demo is scheduled, eliminate most of the false-economy decisions that teams make. The remaining decision — which specific vendor, on which specific tier — is then a tractable shortlist exercise rather than an open-ended landscape sweep.

9.2 Use-case → architecture mapping

The right vendor depends on the right architecture, which depends on the workload. The mapping from workload shape to architecture (and therefore to vendor tier) is reasonably stable.

High variability, high stakes. Healthcare prior-authorization, legal discovery, claims adjudication, KYC review on documents from heterogeneous sources. The 2026 default is Pattern C (hybrid CV+VLM) with explicit HITL gates. Specialization beats generality where the stakes justify the engineering, and the public benchmarks make the gap concrete: Reducto-style hybrid architectures sit roughly 20 percentage points ahead of single-pass systems on adversarial tables, and the gap is wider on documents with mixed modalities (handwriting + print + tables on the same page).

High variability, lower stakes. Customer onboarding with mixed-vendor identity documents, invoice processing across many small vendors, contract abstraction for moderate-stakes commercial agreements. Pattern B (agentic loop) is the right choice here — LlamaParse Agentic, LandingAI ADE on the Team tier, Mistral OCR 3 in its agentic configuration. The reflection loop catches the errors that matter; the absence of specialized routing keeps cost and engineering modest.

Low variability, any stakes. Single-vendor invoices at scale, standardized government forms, regulated forms with strict layouts that change once a year. Pattern A (single-pass VLM) with strong schema validation, or classic template-driven IDP, often wins here on TCO. Adding agentic complexity to a workload that does not need it is paying for variability you do not have.

Long-form reasoning. Financial filing analysis, multi-document loan packets, scientific paper synthesis, e-discovery on document collections. This is where the DABstep result matters: extraction alone is not enough, and even perfect extraction does not yield correct downstream answers without a separate reasoning layer. The right architecture is Pattern B for the parsing layer + a dedicated reasoning agent that operates on the parsed output. Selecting only a parsing vendor for a long-form reasoning use case is the most common shape of “we picked the wrong tool” complaint in 2026.

9.3 The wrong question to start with

Most procurement processes begin with the wrong question: which vendor has the highest benchmark score? This question feels rigorous and is mostly counterproductive.

The reason it is counterproductive: benchmark leaders frequently produce worse results than benchmark followers on specific workloads. ParseBench measures certain things; OmniDocBench measures certain things; RD-TableBench measures certain things. None of them measures what your documents demand. A vendor that scores 84.9% on ParseBench may score 71% on your golden set; a vendor that scores 79% on ParseBench may score 88% on yours. The benchmark rank is one signal; it is not the signal that predicts your production behavior.

The right starting question is sharper: “What is my error budget per output, and what does each error cost me downstream?”

This question is answerable, even before vendor evaluation begins. A healthcare prior-auth shop knows roughly what a wrong decision costs (denial → appeal → potential reversal → reputation cost, often hundreds to thousands of dollars amortized). A KYC team knows what a wrongly approved high-risk account costs (regulatory exposure, sometimes seven figures in fines). A scanning-and-archive workflow knows that an error in a single page costs near zero except for the occasional user complaint when something turns out to be unsearchable.

Once you have the error-cost number, vendor selection collapses to a single objective: minimize cost-per-correct-extraction-with-provenance. The provenance requirement is non-negotiable for any use case that will face an audit; the “correct” definition is workload-specific and is captured by your golden set; the cost is straightforward arithmetic given the per-page parsing cost, error rate, and HITL escalation rate.

This reframing turns vendor selection from a comparison-shopping exercise into an optimization problem. The optimization can be done on a golden set in a week. The comparison-shopping exercise can absorb months without producing a defensible answer.

9.4 ROI per vertical

Specific verticals have specific shapes that worth treating concretely. The four below are the ones where the 2026 economics are clearest.

9.4.1 Healthcare prior-auth

This is the strongest 2026 ROI case in the agentic OCR landscape. The CMS Interoperability and Prior Authorization Final Rule (Centers for Medicare and Medicaid Services 2024), with public-reporting requirements effective March 31, 2026, made manual prior-authorization workflows untenable for health plans operating at any meaningful scale. Plans must now publicly report turnaround times, denial rates, appeal rates, and overturn rates — and the political and reputational pressure of those numbers being public has accelerated agentic OCR adoption faster than any other vertical force.

The public data point is striking. Anterior AI’s workflow on Reducto’s parsing achieved 99.24% accuracy on a prior-authorization golden set (Reducto AI 2026). The deployment was profitable at any commercial-tier parsing price point because the alternative — manual review at roughly 30 minutes per case — was so expensive on a per-correct-decision basis that even premium parsing fees were a rounding error in the total cost.

The architectural default for healthcare prior-auth is Pattern C with explicit HITL gates. Vendor candidates: Reducto, LandingAI ADE (Team or Enterprise plan with HIPAA), LlamaParse Agentic Plus, Mistral Document AI self-hosted for data-sovereignty cases. The decision among these vendors comes down to integration story, HITL workflow quality, and how well the vendor’s golden set predicts behavior on your specific clinical-document types. All four are credible; none is uniformly best.

9.4.3 Financial services — KYC, loan packets, filings

Three sub-shapes within financial services.

KYC document review (identity documents, proof-of-address, source-of-funds documentation) is well-suited to Pattern B with strong schema validation. Vendors that perform consistently here include Reducto (for the document-quality requirements of regulated KYC) and LlamaParse Agentic (for the schema flexibility across document types). The economic case is comfortably above the 25¢ threshold — a single wrongly-approved high-risk customer can generate enforcement risk that dwarfs years of parsing fees.

Loan packets — multi-document submissions for mortgages, commercial loans, asset financing — are the DABstep-style challenge in operational form. The packets contain heterogeneous documents (tax returns, bank statements, employment letters, asset valuations) with cross-document dependencies (the mortgage application references the appraisal; the appraisal references the property tax records; the tax records reference the income statement). Extraction alone is necessary but not sufficient; the system also needs to reason across documents. The right architecture is Pattern B for parsing each document plus a dedicated reasoning layer for cross-document consistency checks. Trying to do this in a single agentic OCR system over-loads the loop and underperforms.

10-K and other financial filings are dense, table-heavy, footnote-laden documents where LlamaParse Agentic Plus is purpose-built for the workload. The premium tier’s pricing (~$0.056 per page) is justified by the structural complexity; in normal Agentic-tier mode, accuracy on filings is meaningfully worse. This is one of the few cases where “pay the premium tier” is the right answer rather than a vendor up-sell.

9.4.4 Customer onboarding (generic)

Generic customer onboarding — account creation, ID verification for moderate-risk products, address verification, basic eligibility — is high-volume, low-to-moderate-stakes, latency-sensitive. Cost-per-page matters because volumes are large; latency matters because users are waiting; accuracy matters but with looser tolerance than KYC.

Mistral OCR 3 at $1–$2 per 1,000 pages is hard to beat for the parsing layer. NuExtract (or NuExtract3) for downstream structured extraction gives you a cheaper, more flexible pipeline than the bundled commercial offerings. The architectural pattern is usually A or B; Pattern C is rarely justified at this end of the stakes/volume curve.

A useful generalization: the higher-volume / lower-stakes end of the workload spectrum is where the AI-native commercial IDP entrants (Nanonets, Docsumo) and the cheap end of the AI-native specialists (Mistral, LlamaParse Cost-Effective) tend to win. The vertical-specific premium vendors (Reducto, LandingAI Enterprise) earn their margin on the high-stakes end.

9.5 When the right answer is “don’t build this”

Most chapters in this book assume you have a use case that justifies agentic OCR and the question is which system to deploy. This section is for the reader who has not yet answered the prior question: should I deploy agentic OCR at all?

The honest answer, in 2026, is “frequently no.” The category is real, the technology works, and there are plenty of workloads where it is the right call — but there are at least as many workloads where building or buying it is the most expensive mistake the team can make. The point of this section is to give you a concrete economic and architectural threshold so you can rule yourself out cheaply, before you have spent a quarter on a pilot that was never going to break even.

9.5.1 Start with the per-page math

Realistic all-in cost-per-correct-extraction for true agentic OCR in mid-2026 breaks down as follows.

Cost component Range per page Notes
Premium agentic parsing (LlamaParse Agentic Plus, ADE Visionary tier) $0.050 – $0.060 The expensive end
Standard agentic parsing (LlamaParse Agentic, ParseBench top entry) $0.010 – $0.020 The 2026 sweet spot
Hybrid CV+VLM (Reducto, LandingAI ADE Team) $0.005 – $0.020
Cheapest “agentic-flavored” parser (Mistral OCR 3 Batch) $0.001 The floor, but closer to single-pass-with-validation than true agentic
Self-hosted open-source (OlmOCR-2 on H100) ~$0.00018 Requires the team to operate it
HITL amortized at typical 5% escalation rate $0.05 – $0.25 Often the largest line item
Eval infrastructure, vendor management, observability $0.01 – $0.10 Hidden but real

Total all-in for a true agentic OCR deployment runs $0.05 – $0.30 per correctly-extracted page. That number is the one you should hold against your downstream value per page.

9.5.2 The one-cent threshold

If your downstream business value is around one cent per page — for instance, a workflow where the entire economic gain from extracting a page correctly is on the order of a penny — agentic OCR is structurally infeasible.

The math:

  • A standard agentic parser at $0.012/page consumes 1.2× your entire per-page ROI on parsing alone.
  • A premium tier at $0.056/page consumes 5.6× the ROI.
  • Even the cheapest credible parser (Mistral OCR 3 Batch at $0.001/page) eats roughly 10% of the per-page ROI before you have paid for engineering, HITL, or anything else.
  • The only parsing-cost configuration that fits is self-hosted open-source. But the open-source path requires an engineering team to operate model serving, evaluation, HITL queues, observability, and vendor-equivalent infrastructure. That team’s loaded cost has to be amortized across enough pages to pay for itself — typically meaning you need millions of pages per month before self-hosted open-source pays for itself against a $0.01/page ceiling.

The honest framing: at $0.01/page downstream ROI, agentic OCR has already eaten your budget before you have written a single line of integration code. The category is wrong for the use case, regardless of how impressive the demos are.

A useful rule of thumb: agentic OCR is economically defensible when your downstream per-page value is roughly 5 cents or higher, and clearly defensible at 25 cents and above. Healthcare prior-authorization decisions, contract clause extraction, KYC document review, and clinical-note coding all sit comfortably above the 25-cent threshold. Commodity invoice line-item extraction, receipt digitization for analytics, and bulk archive indexing typically do not.

9.5.3 When you should not build agentic OCR

Beyond the per-page math, several other workload shapes argue against agentic OCR even when the budget exists:

  • Your real bottleneck is downstream, not parsing. If the documents are native digital PDFs or HTML, classical PDF text extraction plus light cleanup often produces output that is functionally identical to agentic-OCR output. The DABstep result is the reminder that extraction is not reasoning — a system that retrieves answers poorly will not be saved by better parsing, and “we need better OCR” is often a $100k–$500k diagnostic mistake. Chapter 12 covers this in depth; the short version is measure where your chain is failing before you upgrade the parsing layer.
  • Your workload is narrow and stable. One vendor template, three field types, layouts that change once a year. This is the workload that template-driven IDP was designed for, and it is still the cheapest answer on a TCO basis. Adding agentic OCR to a stable narrow workload is paying for variability you do not have.
  • Your stakes are low and your volume is enormous. Hundreds of millions of commodity pages per year, where a 3% error rate is tolerable because the downstream cost of an error is pennies. Hyperscaler basic OCR plus a thin cleanup layer is roughly 50× cheaper for the same downstream outcome. The shift-left thesis (Chapter 6) does not apply at this end of the volume curve.
  • Your latency budget is sub-second. Multi-pass reflection adds wall-clock time. If users expect interactive responses on documents — a chat interface over a PDF, a real-time form-fill assistant — Pattern A (single-pass) is the only architecture that fits, and a single-pass system is by definition not agentic OCR.
  • Your compliance posture is satisfied by a simpler system. Some regulated workflows specifically require deterministic, auditable extraction — not “model with reflection” but “this template, this field, every time.” Agentic OCR’s flexibility can be a liability where determinism is the requirement.
  • You have not built a golden set. Without a golden set against which to measure, you cannot tell whether agentic OCR is improving your accuracy or hurting it. The first investment is the golden set (Chapter 11), not the parser. Teams that skip this step usually deploy the wrong vendor and discover the mistake six months later.

9.5.4 “Don’t build for the sake of it”

The dominant 2026 anti-pattern at the procurement level — distinct from the per-vendor wrapper-washing anti-pattern in Chapter 13 — is building agentic OCR because the field is hot, not because the use case demands it. The pattern shows up as a pilot whose business case was assembled after the vendor selection, with ROI projections that conveniently round up. The pilot ships, the team learns the technology, the ROI does not materialize, the pilot quietly winds down, and the budget is consumed.

The way to avoid this trap is to invert the order. Establish the per-page downstream value first, with a real number that survives stakeholder scrutiny. Establish the all-in cost-per-correct-extraction second, using the table at the top of this section as the floor. Only if the value clears the cost — by at least 3× to give room for the inevitable scope creep — does the question “which vendor?” become live. If the gap is closer than 3×, the right answer is usually to wait. The pricing of agentic OCR is dropping faster than any other component of the modern data stack; today’s marginal workload becomes next year’s comfortable margin.

9.5.5 What to do instead

If the math says agentic OCR is not for you, the alternatives in roughly increasing order of capability are:

  • Native digital text extraction (PDF text layer, HTML parser) for clean documents.
  • Classic OCR (Tesseract, PaddleOCR base) plus regex/template extraction for narrow stable workloads.
  • Template-driven IDP (ABBYY, Hyperscience) when the workload is narrow but the volume justifies a vendor.
  • Hyperscaler OCR APIs (Textract, Document AI, Azure DI) for moderate-volume, moderate-variability workloads where the cost-per-page is genuinely below 0.5 cents.
  • Single-pass VLM (any frontier model with a structured-output prompt) when you need some flexibility but not multi-pass orchestration.
  • Agentic OCR when the use case clears the cost threshold and the documents have the structural complexity to justify the loop.

The right answer is the cheapest one on this list that meets your accuracy and compliance requirements. There is no prize for using the most sophisticated technology; there is a meaningful penalty for using technology your use case cannot pay for.

9.6 The buy-vs-build question

For teams whose workload clears the “should we build this at all” filter, the next question is buy versus build. The framing that has held up in 2026: buy the parser, build the eval.

The case for buying the parser is straightforward. Parsing models are commodities now in the sense that any of a dozen vendors (or a dozen open-source models) can do credible work. The engineering effort to build a competitive parser from scratch — collecting training data, training a vision-language model, building the inference infrastructure — is multiples of the cost of using a vendor for the next three to five years. Almost no team’s competitive advantage lies in being slightly better at parsing than the vendor frontier.

The case for buying surrounding infrastructure (HITL queues, eval tooling, observability) is weaker but still usually positive. Vendors that have been building HITL workflows for fifteen years have engineered things that you would otherwise rediscover. The classic IDP vendors (Hyperscience especially) are often underrated specifically because their HITL infrastructure is genuinely best-in-class, even when their model layer is behind the AI-native specialists.

The case for building your own eval is overwhelming. Only you know what “correct” means for your documents. No vendor’s golden set matches your production distribution. No vendor’s evaluation rubric captures the specific failure modes that matter for your downstream consumers. The vendors that sell pre-built evaluation packages are selling you something that looks like an answer to the eval question but is not — it answers the general eval question, not the yours eval question. Always build your own.

The cases where the build option for the parser becomes correct rather than just defensible: data-sovereignty hard-requirements that no vendor can satisfy, scale at which API spend exceeds engineering loaded cost (typically several million pages per month), and audit requirements that demand a degree of model transparency that no commercial vendor provides. These are real but rare; most teams that think they have these requirements turn out, on close examination, to have softer versions that vendors can satisfy.

A common middle path: use a vendor for the first 12–18 months while the team learns the workload and builds the eval discipline; migrate parsing to open-source self-hosted once the team has the operational maturity to operate it. This pattern produces fewer disasters than either of the extremes (build everything from day one, or buy a vendor and never invest in operational depth).

9.7 A shortlist methodology

The methodology that consistently produces good vendor decisions in 2026 has four steps and takes one to two weeks per shortlist round.

Step 1: Build (or refresh) the golden set. 200–500 hand-labeled documents from your actual production traffic. Include the edge cases your team complains about most. Label them with the schema your downstream consumers actually use. This is the most expensive step and the highest-leverage one; teams that skip it produce vendor decisions that look defensible on paper and fail in production.

Step 2: Pick three vendors across three tiers. One AI-native specialist (LlamaParse, Reducto, ADE, Mistral). One hyperscaler (Document AI, Textract, Azure DI) — even if you do not think you will pick one, the hyperscaler baseline tells you what’s free or near-free in your existing cloud stack. One classic IDP or open-source option (Hyperscience or Granite-Docling, depending on whether HITL or self-hosting is the priority). Three vendors is the right number — fewer leaves you over-anchored on the first; more dilutes the comparison.

Step 3: Run all three on the golden set. Score on the metric that matches your downstream cost (cost-per-correct-extraction-with-provenance is the default; field-level F1 broken down by field is the diagnostic). Track the three numbers that matter: accuracy, cost, escalation rate. Latency too, if it is a binding constraint. Do not score on benchmark scores; you are running your own benchmark now, and yours is the only one that matters for this decision.

Step 4: Decide on data, not demos. The vendor that wins is the one that produces the lowest cost-per-correct-extraction-with-provenance on your golden set, subject to compliance and integration constraints. Resist the pull of the vendor whose demo was most impressive; demos are calibrated to make vendors look good, and the calibration is not your workload.

A separate methodological note: the shortlist should be re-run annually. The vendor landscape changes faster than procurement cycles. A vendor that was the right choice last year may not be this year; a vendor that was not credible last year may have caught up. Annual re-evaluation is good hygiene and is cheaper than the alternative (sticking with a vendor whose relative position has deteriorated).

9.8 A walked example: 80k pages/month insurance underwriting

To make the framework concrete: an insurance team running underwriting at 80,000 pages per month, mixed digital and scanned PDFs, strict compliance requirements (SOC 2 Type II, HIPAA BAA), three-month deployment window.

Applying the framework:

Volume: medium (80k pages/month). Rules out the very-low-volume “just use a hyperscaler” answer; rules out the very-high-volume “self-host open-source” answer (the team is not yet operating at the scale where self-hosted economics are decisive). Sits squarely in the AI-native-specialist sweet spot.

Variability: wide. Mixed insurance carriers, mixed document types within each carrier’s submission. Rules out template-driven IDP as the primary solution.

Compliance: SOC 2 + HIPAA. All four AI-native specialists clear SOC 2; only LandingAI ADE and Reducto have published HIPAA BAA options at the time of this writing. (LlamaParse and Mistral can typically negotiate HIPAA on enterprise plans but it is not a published default.) The compliance filter narrows the shortlist to ADE and Reducto, plus one hyperscaler or classic IDP as a sanity-check baseline.

Self-host vs. SaaS: SaaS is acceptable here given the BAA and ZDR options. Self-hosting is not required.

Stakes: moderate-to-high. Wrong underwriting decisions cost five to six figures per incident; volume-times-error-rate produces meaningful expected loss. Justifies Pattern C investment if the parsing layer cannot meet accuracy targets in Pattern B.

The shortlist that emerges: Reducto, LandingAI ADE Team plan, and AWS Textract (as the hyperscaler baseline). Run all three on a 300-document golden set of representative underwriting submissions. Score by cost-per-correct-extraction-with-provenance plus HITL escalation rate. Whichever wins on those metrics — and the answer is genuinely not obvious in advance for this workload — gets deployed.

This shape of decision, walked through with explicit numbers from a real-ish use case, is what good vendor selection looks like. The math is straightforward; the discipline is the rare part.

TipNamed takes

Most teams pick a vendor before they pick a metric. This is backwards. Choose the metric that determines your downstream cost, then choose the vendor that minimizes that metric.

Build-vs-buy is the wrong frame. Buy the parser, build the eval. Reuse vendor parsers; never reuse vendor evals.

The shortlist methodology — three vendors, one golden set, one cost-per-correct-extraction number — is unglamorous and consistently produces better decisions than any amount of vendor-deck reading.