The Executive Playbook for the Agent-to-Agent Economy
Every roadmap answers one question, whether it says so or not: who acts. This book is about the line moving right now — the Mandate Boundary — separating what software may decide alone from what still needs a human's signature, and what a product leader should build before that line has finished moving.
Get the full PDF on Ko-fi →When Software Hires Software — an executive playbook on the agent-to-agent economy: where the Mandate Boundary sits today across procurement, banking, mortgage lending, and healthcare, and what to build before it finishes moving.
Copyright © 2026 Anuj Sadani.
All rights reserved. Quote it in a review or a piece of criticism with attribution; anything further, please ask.
First edition · 2026
Every roadmap answers one question, whether it says so or not: who acts. For thirty years the answer was a person, and the software just carried the paperwork. That is the assumption breaking right now, and it is breaking on the buying side first.
Look at what actually shipped between April 2025 and April 2026, not what got keynoted. Mastercard built Agent Pay, tokens that bind a card to a specific agent, a specific merchant, and a specific consent policy 1. Visa built the Trusted Agent Protocol, a verified identity for an agent plus a signed consent record from the issuer 1. Google built AP2, which turns a purchase into three cryptographically signed contracts — an Intent Mandate for what you asked for, a Cart Mandate for what got assembled, a Payment Mandate for what gets charged 2. None of that is a chatbot feature. It is financial infrastructure for a counterparty that isn't a person, built by the companies whose entire business is knowing who is allowed to move money.
That is the fact this book is built around: the frontier of AI product work has quietly moved from can software understand something to can software be trusted to act on it, alone, with money and consequences attached. Call it what it is. Software is starting to hire software — to shop from it, negotiate with it, pay it, and increasingly, to answer to it when something goes wrong. Every domain chapter that follows is one industry's version of this same shift. This chapter is the shift itself, and what a leader does about it before the shape of 2030 is remotely settled.
Forget the debate about whether AI is "agentic" — that word is mostly marketing residue at this point (Gartner counted about 130 vendors with real agentic capability out of thousands claiming it, and predicts over 40% of agentic AI projects will be scrapped by the end of 2027 for cost, unclear value, or missing risk controls 3). The debate worth having is narrower and more useful: for a given decision, where does authority to act without a human in the loop currently sit, and how far are you willing to move it?
Call that the Mandate Boundary — the line separating decisions a system may execute unaccompanied from decisions that still require a human counter-signature. Every mandate protocol built in the last eighteen months is, underneath the branding, a machine for drawing this line precisely and making the drawing enforceable: AP2's Intent and Cart Mandates exist specifically to separate "what I authorized" from "what the agent decided" so the boundary survives a dispute 2. Visa's Trusted Agent Protocol exists to prove an agent's identity to a merchant before the boundary gets tested 1.
Your roadmap is not really a list of features. It is a series of bets about where you move the Mandate Boundary, for which decisions, on what evidence, and how fast you can move it back if the evidence turns bad. Once you see the roadmap this way, "what should we build for 2027 versus 2030" stops being a features question and becomes a governance question — which is a much more answerable one, because governance questions have precedent even when the technology doesn't.
The protocol for letting software transact is now mostly solved. What remains unsolved — and what every domain in this book is really fighting over — is who is accountable when the transaction is wrong.
Here is the uncomfortable part for anyone used to eighteen-month planning cycles: the frontier moved and reversed inside a single one. OpenAI launched in-chat Instant Checkout on the Agentic Commerce Protocol in September 2025, built with Stripe, aimed squarely at the "customer never leaves the chat" vision every 2025 keynote gestured at. By March 2026 — six months later — OpenAI retired it. Reported adoption topped out somewhere between "about a dozen" and "closer to thirty" live merchants, and Walmart found checkout inside ChatGPT converting roughly three times worse than a click-through to its own site 4. The pattern that survived wasn't "AI closes the sale." It was "AI helps you discover, and you still buy where you already trust the receipt."
That is not a story about a bad product. It is a story about the actual shape of this wave: pieces of it move immediately and permanently — protocol governance for MCP and A2A both landed with the Linux Foundation inside about a year of each launching 5 6 — while other pieces that looked equally inevitable in a demo (consumers letting an agent complete the whole purchase) stall on trust that hasn't been earned yet, and might not be earned on anyone's preferred timeline. Both things are true about the same eighteen months. A roadmap built on either "none of this is real yet" or "all of this is inevitable by next quarter" will be wrong regardless of which one you pick, because the wave doesn't move as one wave. It moves as several waves at different speeds, and the speed each one travels at is set by how expensive it is to be wrong, not by how good the technology has gotten.
That single fact is the whole argument for building flexibility into a 2030 plan instead of a fixed destination. Nobody who published a roadmap in early 2025 predicted that the most hyped consumer agentic-commerce feature of the year would be dead by its own six-month anniversary while the identity and payment rails underneath it kept right on standardizing. The rails were boring, so they shipped. The consumer-facing bet was exciting, so it got canceled. Build for the second pattern to repeat, because it will keep repeating: the plumbing keeps compounding quietly, the exciting front-end bets keep dying in public.
Here is the part that should reorganize how you think about differentiation. MCP and A2A are not proprietary technology anymore — they are Linux Foundation projects, governed the way Kubernetes is governed, built specifically so no single vendor can own them 5 6. If your 2027 plan is "we will expose an agent interface," you have described table stakes, not a strategy. Every competitor with an engineering team can expose the same interface by next quarter, and several already have.
The moat was never going to be in the interface. It has to be in what stands behind it when an agent — yours or someone else's — asks a hard question and needs an answer a court, a regulator, or an auditor will accept later. Three things sit behind a durable answer, and none of them are the protocol:
Verified facts with provenance. Not "the model's best guess," but a value with an evidence trail — which document, which page, which rule version, computed when, by what — attached to it permanently. This is expensive to build and boring to describe, which is exactly why it's defensible.
Warranted outcomes. A price attached to being right, not just to answering. The difference between a vendor who tells you a decision and a vendor who stakes something on that decision being correct is the difference between a feature and a business.
Accountable judgment. A specific party — a licensed underwriter, a named institution, a warranty-backed system — who absorbs the consequence when the decision is wrong. Software can execute a mandate. It cannot yet be the thing a regulator holds responsible when the mandate was a mistake, and nothing in any protocol shipped so far changes that.
Every domain chapter that follows names the version of these three things that applies to one industry, because "verified facts, warranted outcomes, accountable judgment" is not a slogan — it is the actual checklist for telling a real product from a demo in this wave. If your 2027 plan doesn't strengthen at least one of the three, it is decoration on top of someone else's infrastructure, and it will feel that way to your customers within a year.
Naming the moat is only half the roadmap. The other half is a pricing decision every one of these initiatives eventually forces, and getting it wrong is a common way to build the right infrastructure and still lose the business. Four patterns recur across every domain in this book, and the durable products almost always combine two of them rather than picking just one.
Metered access prices the thing closest to a commodity — a single tool call, a single check, a single page processed — by the unit, the way a document-extraction API bills per page or per credit. This is the right model for atomic, deterministic, low-stakes operations: cheap enough that volume is the business, and simple enough that a buyer can predict their bill before they commit to it. It is also the model most exposed to commoditization, because a metered price is trivially comparable to a competitor's metered price. Use it for the parts of your product that really are commodities, and stop pretending otherwise.
Outcome-warranted pricing charges for being right, not for answering — a fee per approved authorization, per closed loan, per resolved case, with an actual guarantee or fee-at-risk attached if the outcome turns out wrong. This is the model that only becomes available once you've built the "warranted outcomes" leg of the moat described above, and it is worth far more per transaction than metered access, because the buyer is no longer purchasing an answer. They're purchasing the transfer of a risk they used to hold themselves.
Tiered subscription with role-gated access sells a seat, not a call — a monthly or annual fee that unlocks a specific bundle of tools, data freshness, and volume, with the tier determining not just price but rate limits and which tools are even visible to that user. This is the commercial expression of the same progressive-disclosure logic that keeps a well-designed agent's tool count manageable: a junior role's seat exposes a handful of tools at a modest rate limit; a senior role's seat exposes more tools, deeper data, and a higher ceiling; an enterprise tier adds custom overlays and effectively uncapped throughput. Price the tier to the value of what a role can see and do, not to how much compute it happens to consume.
One-time integration plus recurring usage front-loads a fee for the unglamorous work of connecting to somebody else's system of record — a lender's loan-origination system, a payer's claims platform, a hospital's EHR — and then layers a recurring fee, usually one of the first three patterns, on top once the connection is live. Skip this and you'll underprice the hardest, least repeatable part of the sale: getting into the account at all.
Every domain chapter that follows names its own version of this menu, because the question "what do we charge for" deserves the same concrete answer as "what do we build," and a roadmap that only answers the second question will spend 2027 giving away, metered by the call, the exact thing it should have been warrantying by the outcome.
Treat 2027 and 2030 as two different kinds of bet, not two points on the same line.
2027 is a build horizon. The pattern already visible across every domain in this book is software recommends, a human authorizes — a mandate gets assembled with evidence, a specific accountable person signs it, and the system's job is to make that signature fast, well-informed, and cheap to defend later. This is buildable today with technology that already exists: evidence-bundled decisions, role-scoped tool access so a human reviewer sees exactly what they need and no more, and outcome-priced pilots that put your own money where your accuracy claims are. Build this now. It compounds — every decision it makes correctly becomes training data, a defensible track record, and one more data point in the warranty argument you'll need to expand the Mandate Boundary later.
2030 is a positioning horizon, not a build target. The pattern to prepare for, without betting the company on its exact timing, is software authorizes within a policy a human already set — continuous, standing mandates re-evaluated as facts change, rather than one-time approvals re-litigated from scratch every time. Mortgage lending's "always pre-approved" is a plausible version of this. Auto-locking a rate and closing a loan without a human touching it is not plausible on any near-term timeline, for the same reason full agentic checkout stalled: the cost of being wrong is high and the trust to be delegated hasn't been built yet. The organizations that win this horizon won't be the ones who guessed the date correctly. They'll be the ones who spent 2027–2029 accumulating the evidence trail, the warranty capacity, and the regulatory relationships that let them move the Mandate Boundary the moment the trust actually exists — while their competitors are still arguing about whether it exists yet.
Practically, that means every initiative on your roadmap should answer two questions before it gets funded. First: which side of the Mandate Boundary does this sit on today, and what evidence would move it? Second: if the specific bet inside this initiative turns out to be six-month-hype rather than durable infrastructure — the Instant Checkout pattern rather than the AP2 pattern — how much of what we built survives the pivot? Facts infrastructure, evidence schemas, and accountability relationships survive almost any pivot. Bets on a specific vendor protocol or a specific consumer behavior do not. Weight your 2027 spending accordingly, and you get the flexibility to pivot without having wasted the year finding out you needed it.
The rest of this book walks that same S-curve through four places where the money is already real, not speculative: procurement, where agents are starting to shop from other agents; banking, where the wallet itself is becoming an algorithm; mortgage lending, where a decision infrastructure is quietly replacing what used to be a document pipeline; and healthcare, where a claim is starting to argue with itself before a human ever sees it. Each chapter names the specific Mandate Boundary in that domain, where it sits today, and what it would take to move it. Start wherever your business already touches one of these four — the mechanics transfer faster than you'd expect, because underneath the industry-specific vocabulary, it's the same question every time: who is allowed to act, on what evidence, and who answers for it when they're wrong.
Somewhere in your target account list right now, a piece of software is deciding whether your company is worth a human's time — and it is making that call faster than a person could finish loading your homepage. It does not watch your demo video. It does not read your case studies for tone. It checks whether your claims can be verified without picking up a phone, and if they can't, it drops you from the shortlist before anyone on your sales team learns you were ever being considered.
Call that check the Vendor Handshake: the exchange of signed, checkable claims — specs, security attestations, pricing, compliance status — that a buying agent runs against a vendor before it will do anything a human buyer used to do first, like shortlisting, comparing, or requesting a quote. Every domain in this book has its own version of the moment one piece of software decides whether to trust another before a person gets pulled in. In procurement, that moment is already happening, in production, at scale — just not in the form most vendor marketing sites were built to survive.
This chapter answers a narrower question than "will AI change B2B sales." It answers: when the buyer, the comparison-shopper, and increasingly the person drafting the RFP response are all software, what should a vendor build in the next twelve months, and what should it merely position for?
Start with what's actually been funded, because procurement is one of the few corners of this wave where the capital allocation tells you more than the keynotes do.
Profound, a company that exists to make brands legible to AI answer engines rather than search engines, raised a $96 million Series C at a $1 billion valuation in February 2026, becoming the category's first unicorn. Seven months later, in September 2026, it raised a $180 million Series D at $1.8 billion, co-led by Sequoia and Kleiner Perkins, with revenue up roughly 3x over that same six-month stretch and more than 1,000 enterprise customers on the books, including Comcast, Estée Lauder, and Walmart 7. That is not hype-cycle money. That is a market that has already decided generative-engine optimization — GEO, the practice of making your company's facts legible to the models doing the summarizing — is worth real budget, twice, in under a year.
Incumbents are reading the same signal. Adobe agreed to buy Semrush, the SEO analytics platform, for $1.9 billion in November 2025, explicitly to fold GEO capability into Adobe Experience Cloud before someone else owns the category 8. When a company with Adobe's balance sheet pays nearly $2 billion for a bolt-on rather than building it, that's a company telling you the build-versus-buy math already favors moving fast.
And the demand side backs this up with actual commerce data, not just vendor claims. On Shopify's Q1 2026 earnings call, President Harley Finkelstein reported that AI-referred traffic to Shopify stores was up roughly 8x year over year, with orders from AI-powered search up nearly 13x — new buyers arriving through AI channels at close to twice the rate of other channels 9. By the Q2 2026 call in August, both AI traffic and AI-attributed orders had tripled year over year again, and Finkelstein noted that AI search running on Shopify's own structured catalog data converted at roughly twice the rate of AI search relying on scraped product pages 9. Read that last detail twice — it's the whole chapter in one sentence. The advantage wasn't "we're in more AI answers." It was "our facts are structured well enough for the answer engine to trust them without scraping."
Meanwhile, distribution is fragmenting into open warfare. Amazon expanded its robots.txt blocklist to keep OpenAI's ChatGPT-User and OAI-SearchBot crawlers off its product pages in late 2025, protecting an advertising business that depends on shoppers browsing Amazon directly rather than through a chat window that skips the ads 10. Then Amazon sued Perplexity, accusing its Comet browser of disguising its shopping agent as a regular human session to keep scraping after being blocked — and won a preliminary injunction in March 2026, before the Ninth Circuit vacated it in August on the technical grounds that it was the user, not Perplexity, accessing Amazon's servers under federal computer-access law 10. Nobody in that fight was arguing about whether agents will shop. They were arguing about who gets paid when they do.
Set all of that against the thing that was supposed to be the headline feature: OpenAI's Instant Checkout, launched in September 2025 to let a purchase complete entirely inside a chat window, retired roughly six months later after adoption topped out somewhere between a dozen and thirty live merchants and Walmart found in-chat checkout converting at about a third the rate of a click-through to its own site 4. Consumers, it turns out, still want to land on a page they recognize before money moves. That is the same trust gap this book's opening chapter identifies across every domain — protocol infrastructure compounds quietly while the consumer-facing bet dies in public.
So here is the honest 2027 map for procurement, and it is not the map most GTM decks are drawing. The wedge isn't consumer checkout autonomy — that stalled, and there's no evidence it's coming back on anyone's preferred timeline. The wedge is being legible and verifiable to the agents doing the comparing, whether those agents are consumer-facing answer engines or the buyer-side procurement agents inside your prospects' companies. GEO money is real because it answers a question every vendor already has: does an AI system, asked about my category, say my name accurately, or does it hallucinate my competitor's pricing into my slot. Verified-claims infrastructure is the underbuilt half of that same question: once you're mentioned, can a buying agent check what was said about you without a phone call — and will it trust the answer enough to shortlist you.
Here is the mechanic underneath "software hires software," made concrete for procurement.
A human buyer on a vendor shortlist tolerates ambiguity because a human can ask a follow-up question, read tone, and build trust across a sales cycle. A buying agent doesn't get that luxury, and — more importantly — it isn't supposed to. Its whole value proposition to the human who deployed it is that it can shortlist, compare, and draft a recommendation without a six-week discovery process. That only works if the facts it's comparing are structured well enough to be checked mechanically rather than taken on faith. This is the Vendor Handshake in practice, and it has a fairly specific shopping list:
Signed specifications, not marketing copy — a machine-readable data sheet with a cryptographic signature tying the claim to the entity that made it and the date it was made, the B2B equivalent of the structured catalog data Shopify's own numbers show converting twice as well as scraped pages 9.
Security and compliance attestations exposed at a well-known, machine-queryable endpoint rather than buried in a PDF a salesperson emails after a discovery call — SOC 2 status, data residency, uptime history, the answers that used to live inside a security questionnaire a human filled out under deadline pressure.
Pricing and contract terms a buying agent can evaluate against its mandate without a call — not necessarily public list pricing, but structured enough that an agent operating inside a spend policy can determine fit or disqualification on its own, the same way Google's AP2 turns a consumer purchase into signed Intent, Cart, and Payment mandates precisely so a dispute can be resolved by checking a record instead of re-litigating a conversation 2.
A verification path that doesn't depend on the vendor's own say-so — third-party attestation, audit trail, or a protocol-level signature the buying agent's infrastructure can validate independently, the same trust-before-transaction logic Visa built into the Trusted Agent Protocol so a merchant can confirm an agent's identity before anything moves 1.
None of this requires a proprietary interface. MCP and A2A are Linux Foundation projects now, governed the way Kubernetes is governed, specifically so no single vendor can lock up the pipes 5 6. A vendor exposing an agent-readable endpoint next quarter is describing table stakes, exactly as the opening chapter of this book warns about any "we'll expose an agent interface" roadmap line. The differentiation was never going to be in having a feed. It's in whether anything you put on the feed can survive being checked.
That's the domain-specific version of the Mandate Boundary this book keeps returning to. In procurement, the Mandate Boundary sits at the shortlist and the RFP response: a buying agent may gather, compare, and rank vendor claims unaccompanied, but the moment those claims translate into a signed commitment — a contract, a purchase order, a vendor relationship a company will depend on for years — a human still has to countersign, because nothing about a well-formatted claims feed changes who absorbs the consequence if the claims were wrong.
MCP and A2A being free protocols means the interface was never the business. The three things that were always going to be the business, applied to procurement specifically:
Verified facts with provenance. Not a spec sheet a vendor wrote about itself, but a claim with an evidence trail attached permanently — which security audit, which pricing tier, which compliance certification, issued by whom, verifiable independently of the vendor repeating it. This is what a verified-claims feed actually is, once you strip the GEO branding off it: provenance infrastructure for B2B facts, built the same way facts infrastructure gets built in every other chapter of this book — boring, expensive, and exactly why it's defensible once built.
Warranted outcomes. The uncomfortable question a claims feed has to answer eventually: what does the vendor stake on the claim being accurate? A spec sheet with no consequence attached is marketing. A spec sheet a vendor is willing to warrant — refund the deal, cover a security incident, eat a penalty if the attestation was stale or wrong — is infrastructure a buying agent's operator can actually rely on, because the agent's mandate can now price the risk of being wrong instead of just trusting the claim. This is the same distinction the opening chapter draws between a vendor who tells you a decision and one who has priced being wrong into the product.
Accountable judgment. When a buying agent shortlists a vendor off a bad claim and the resulting purchase goes badly, who answers for it? Not the model — the model can't be fired, sued, or held to a contract. The answer has to be a specific, named party: the vendor that signed the attestation, the procurement lead who countersigned the mandate at the Mandate Boundary, or — increasingly, this is the actual product opportunity — a third-party verification provider whose entire business is standing behind the claim so that both sides of the transaction can point to a name when something breaks. That is a genuinely new role in the B2B stack, and right now almost nobody occupies it cleanly.
A shortlist an agent can generate in ten seconds is worth nothing to a buyer who still has to spend six weeks finding out whether any of it was true.
That's the gap the current wave of GEO tooling doesn't close, because tracking whether you're mentioned accurately is a monitoring product, not a trust product. The vendors who win the next stretch of this build will be the ones who stopped asking "are we visible to the agent" and started asking "can the agent verify us without calling us" — and then went and built the warranty behind the answer.
Run procurement's version of the moat through Chapter 0's menu and four concrete products fall out, not just a direction.
A verified-claims feed — the signed specs, attestations, and pricing this chapter has already described — is naturally a tiered subscription: a base tier covering a fixed claim set and a modest number of agent verification pulls per month, a higher tier unlocking more claim categories (security, compliance, custom pricing) and a higher rate limit on how many times a buying agent can query it, because the vendor's real cost scales with verification volume, not with having claims at all. Layer a metered verification fee on top for agents outside a buyer's existing subscription — a pay-per-check option for the long tail of buying agents a vendor has no standing relationship with.
An agent-readiness audit — bringing a vendor's product feed and well-known endpoints up to a standard a buying agent can actually parse — is a clean one-time integration fee, priced like a security audit, with an optional ongoing monitoring subscription for vendors who want to be re-certified as protocols evolve.
The third-party verification and warranty role this chapter identifies as a genuinely open position in the stack is where outcome-warranted pricing belongs: a per-claim or per-deal verification fee that scales with the value of the transaction it's underwriting, with a real payout if a claim the verifier certified turns out to be false. That is a materially different — and larger — business than selling access to a claims database, for the same reason an insurance premium is worth more than a subscription to weather data.
And the 2030 buyer-agent RFP responder this chapter already prices per qualified opportunity is outcome-warranted pricing taken to its logical endpoint: a vendor pays only when its agent's answers actually advanced a real deal, which is the only pricing model a buyer will accept for a system negotiating on their behalf without supervision.
Here the evidence runs out and the extrapolation starts, so treat what follows as informed positioning, not a forecast with sources behind it.
The plausible next stage, if the Vendor Handshake becomes standard infrastructure the way TLS became standard infrastructure for web traffic, is a buyer-side procurement agent that doesn't just shortlist vendors — it drafts and negotiates the RFP response itself, querying vendor agents directly over A2A for warranted answers to security questionnaires, SLA terms, and pricing tiers, priced by the vendor per qualified opportunity rather than per lead 6. That's a genuinely different shape of GTM motion: instead of a vendor's sales team chasing a signal that a company is "in-market," a vendor's agent is fielding structured, machine-generated RFP queries around the clock and only escalating to a human seller when a deal clears a threshold worth a human's time.
The harder version of this — a standing procurement mandate, re-evaluated continuously as a vendor's certifications, pricing, or security posture changes, the way this book's mortgage chapter describes an "always pre-approved" borrower — is a plausible 2030 pattern for repeat, low-variance B2B purchases: consumables, commodity SaaS seats, routine service renewals. It is not plausible yet for anything with real switching cost or real risk, for the same reason full agentic checkout stalled in retail: the price of being wrong is still higher than the trust that's been built to be delegated. This is my inference from where the infrastructure is heading, not a sourced benchmark — nobody has shipped a standing B2B procurement mandate at scale yet, and the regulatory and contractual questions underneath one (who's bound, under what authority, revocable how fast) are further from settled than the payment rails are.
What's worth building toward now, regardless of exactly when 2030 arrives, is the accountability relationship a standing mandate would require: a named party who can be pointed to when a vendor's claim turns out to be stale or a purchase goes wrong. That relationship is worth building years before the protocol matures, because it's the part that doesn't get thrown away if the specific 2030 pattern turns out wrong and something else wins instead.
Procurement was never really a question of who gets to shop. It was always a question of who gets to spend the company's money and who answers for it after. The next place that question shows up, in a more literal form, is the wallet itself — which is exactly where this book goes next.
A corporate card used to need a person to swipe it. By 2027 it mostly needs a policy instead — and the fight over who writes that policy, not who processes the swipe, is where the actual money in this chapter is.
Call the thing the policy governs the Policy Wallet: not an account a human authorizes one purchase at a time, but a standing set of conditions — vendor, category, price band, cadence, counterparty — that a piece of software checks against every attempted payment and either honors or blocks, with nobody awake to watch it happen. The Policy Wallet is the Mandate Boundary from Chapter 1 poured into a specific, cash-shaped mold. Every mandate protocol shipped in the last eighteen months — Mastercard's Agentic Tokens, Visa's Trusted Agent Protocol, Google's AP2 — is a machine for drawing that boundary around money specifically: what an agent may spend, on what, before a human has to look at it again. The protocols answer how a payment gets authorized. They do not answer what the policy should say, who is on the hook when the policy is wrong, or how a bank proves any of it to a regulator eighteen months later. That gap is the whole chapter.
Start with what's no longer in question, because treating it as an open competitive space is the fastest way to waste a 2027 budget.
Mastercard announced Agent Pay in April 2025 — Agentic Tokens that bind a card to a specific agent, merchant, and consent policy, built on the same Mastercard Digital Enablement Service tokenization that already powers Apple Pay and Google Pay 1. Visa's Trusted Agent Protocol followed that October, giving merchants a way to verify an agent's identity before checkout rather than guess whether it's a bot 1. Both networks have since moved from announcement to production. Mastercard says the first fully autonomous agentic payment — an AI agent booking a ride through a mobility provider called Hoppa, authenticated by DBS Bank with no human confirming at the point of sale — happened in Singapore on March 4, 2026, and in June 2026 the company launched Agent Pay for Machines with more than thirty partners spanning cards, bank accounts, and stablecoins on one settlement layer 11. Visa, meanwhile, struck a partnership with OpenAI the same month to embed tokenized Visa credentials directly into agent-initiated checkout inside ChatGPT and Codex, with spending limits, merchant categories, and required-approval rules enforced at the network level 12. And in May 2026, Google donated AP2 itself to the FIDO Alliance, alongside a Mastercard-built companion standard called Verifiable Intent, with new FIDO working groups on agentic authentication and agentic payments now chaired jointly by Mastercard, Visa, Google, OpenAI, and Amazon 13.
Three competitors who spent decades fighting each other for interchange share converged on overlapping, interoperable standards for machine-initiated payments inside about thirteen months, and then handed the governance of those standards to a neutral standards body so no single one of them could own it outright. That is not a market with room for a new entrant building "a payment rail for agents." That market closed before most product teams finished their first planning cycle on the topic. If a line item on your 2027 roadmap is infrastructure that competes with Agent Pay, TAP, or AP2 directly, kill it now and redeploy the budget.
The opportunity that survives this consolidation sits one layer up, exactly where Chapter 1 said the moat would be once a protocol goes free: the policy semantics layer that decides what a Policy Wallet is actually allowed to do, and the evidence layer that proves who's really on the other end of the transaction. Neither Mastercard nor Visa nor Google is in the business of encoding your specific procurement rules — "this agent may reorder consumables within five percent of contract price, from this vendor list, capped at this monthly ceiling" — into an enforceable Intent Mandate. Somebody has to compile that policy, in that specificity, and somebody has to package the KYC/KYB evidence that proves the counterparty accepting the payment is who and what it claims to be, especially when the counterparty is itself software. That compiling and packaging work is unglamorous, vertical-specific, and exactly the kind of boring infrastructure that compounds instead of getting commoditized by the next protocol release.
It also happens to be the work a CFO actually needs done before any of this touches a real treasury account. No finance leader signs off on a Policy Wallet because the rails are elegant. They sign off when three specific questions have answers they can put in a board deck: what, precisely, is this agent allowed to do without me; what happens to the money if that policy gets exploited or simply misfires; and can I explain the wallet's behavior to an auditor, a regulator, or a plaintiff's lawyer without pointing at a model and shrugging. Those three questions are, not coincidentally, Chapter 1's moat restated as a procurement checklist — evidence, warranty, explainability — and they are the actual gate between "we piloted an agentic payments demo" and "we let it move real money unattended." The rest of this chapter is about what has to exist on each side of that gate.
Here is the part that changes the shape of the problem rather than just its scale. Everything so far describes one agent — a shopping assistant, a procurement bot — transacting against a human-run business. That's the easy half. Mastercard built Agent Pay for Machines specifically for the other half: machine-to-machine commerce, where an IoT sensor pays per API call, a logistics agent pays a warehousing agent for storage by the hour, and neither side of the transaction has a person checking a dashboard 11. When the payer's wallet and the vendor's invoicing system are both agent-operated, "software hires software" stops being a checkout-page feature and becomes the entire commercial relationship.
That reshapes KYC and KYB, because the question is no longer just "is this business real." It's "which specific agent is acting, under whose authority, with what spending limit, and can that be proven independently of the agent's own say-so." An emerging shorthand for this — Know Your Agent — is circulating among identity vendors and standards groups as of 2026: the idea that every acting agent needs to be cryptographically bound to a verified human or business principal, with its permission scope machine-readable and its behavior auditable over time 14. Treat that framing as directional rather than settled; there is no single regulator-blessed KYA standard the way there is a chartered bank behind KYC. But the underlying problem it's naming is real and it's exactly the one a Policy Wallet has to solve before a CFO will let it run unattended: prove not just that the vendor exists, but that the specific piece of software claiming to represent the vendor was actually authorized to send that invoice.
This creates a specific failure mode. In a human-to-business transaction, a confused or malicious counterparty eventually runs into a person who notices something is off — a purchasing manager who flags an unusual invoice, a customer who disputes a charge. In an agent-to-agent transaction, that noticing has to be engineered in deliberately, because by default nobody is watching. Whichever side of the transaction is holding the better evidence trail — the more complete, better-provenanced record of what was authorized, by whom, under what policy — wins the dispute when it eventually happens, whether that dispute plays out as a chargeback, an audit finding, or a lawsuit. That is the KYC/KYB evidence packager business in one sentence: sell the receipts that make a machine-to-machine dispute resolvable by something other than "our agent said your agent said."
Chapter 1's three-part moat — verified facts with provenance, warranted outcomes, accountable judgment — translates into this domain almost without modification, which is a good sign that it's the right framework rather than a slogan.
Verified facts with provenance is literally the KYC/KYB evidence trail, and the infrastructure to generate it is arriving on a hard regulatory clock rather than a product roadmap. Under the EU's eIDAS 2.0 regulation, banks, insurers, and payment providers must accept the EU Digital Identity Wallet as an authentication method for onboarding and strong customer authentication by December 2027, with each member state required to have at least one government-issued wallet live a year earlier 15. That deadline is real, but the rollout underneath it is not going smoothly — reporting as of September 2026 describes the wallet stalling in roughly two dozen member states even as the regulatory clock keeps running — a reason not to bet a 2030 plan on flawless infrastructure arriving on schedule 15. India is running a parallel experiment with a different mechanism: the DPDP Act's Consent Manager framework goes live on November 13, 2026, formalizing the licensed-broker model that the country's Account Aggregator network has been operating since 2021, where a regulated consent manager moves encrypted financial data between institutions without ever holding the decryption key itself 16. As of late 2025, seventeen licensed Account Aggregators connect more than 135 financial information providers, with roughly 38 percent of borrowers reportedly enabled on the network — figures worth treating as directional, since they come from industry secondary sources rather than the regulator itself 17. Both systems are, underneath the acronyms, the same idea: a verifiable, consent-scoped, revocable evidence trail attached to a financial fact, which is precisely what a KYC/KYB evidence packager needs as raw material.
Warranted outcomes is where this domain gets genuinely uncomfortable, because nobody has actually settled who pays when a Policy Wallet gets it wrong. Agentic payments break the assumption baked into decades of consumer-protection law that a transaction has one clean moment of authorization by one identifiable person. Under Regulation E in the US, a transfer is generally presumed authorized once a consumer hands credentials to an agent, even if that agent then acts outside what the consumer actually intended — which means the consumer-protection backstop that would normally absorb a bad agent decision may simply not apply 18. Legal analysis of the current wave of agent payment tools is blunt about the state of play: the question of who absorbs the loss — the person who authorized the agent, the developer who built it, or the merchant who received the funds — will likely be settled by whoever litigates first, and no company wants to be that test case 18. That uncertainty is not a reason to wait. It's the actual product opportunity. A mandate policy engine that merely executes a rule is a feature. A mandate policy engine backed by a warranty — a vendor who prices, underwrites, and stands behind the accuracy of its own policy enforcement — is a business, for the same reason an insurer is a business and a weather app is not.
Accountable judgment has the sharpest edge of the three, because it's not really a business risk here — it's a legal wall. The CFPB's 2022 guidance established that a lender cannot use a credit-decision model too opaque to produce specific, accurate reasons for turning someone down; "the black box told us to" is not, and has never been, an acceptable adverse-action notice under the Equal Credit Opportunity Act and Regulation B 19. In May 2025, the CFPB withdrew that guidance along with dozens of other interpretive documents as part of a broader deregulatory push 19. That withdrawal did not repeal ECOA or Regulation B — the statute is still the statute, and a creditor still has to produce specific, accurate reasons for a denial or face liability under the underlying law, guidance or no guidance. If anything, the withdrawal makes the constraint sharper for a builder, not softer: the safe harbor of clear agency guidance is gone, and what's left is bare statutory exposure that gets tested in court rather than pre-cleared by a regulator. This is exactly why SMB cash-flow underwriting built from consented bank data with reason codes traceable to specific transactions is more defensible than a generic underwriting model wrapped in a chat interface. Explainability here is not a nice-to-have UX layer. It is the difference between a product a bank's counsel will approve and one they legally cannot.
A wallet that can't explain why it said no isn't a product a bank can ship. It's a liability nobody has been sued over yet.
The Policy Wallet's economics split cleanly along Chapter 0's menu, and the split matters because pricing a mandate policy engine like a SaaS seat, instead of like an insurance product, is the fastest way to underprice the actual risk you're taking on.
A mandate policy engine — the compiled, enforceable version of "this agent may reorder consumables within five percent of contract price" — is sold as a tiered subscription per wallet or per policy ruleset, with the tier gating both the complexity of policy a customer can express and the transaction throughput the engine will authorize per hour, the same kind of rate-limited access-tier structure that governs how much a role can do in the mortgage chapter's product ladder. Layer a metered per-transaction fee on top for actual mandate checks executed above the tier's included volume, because the marginal cost of evaluating one more transaction is real and should be priced as such, separately from the fixed cost of having a policy engine at all.
A KYC/KYB evidence packager for agent and non-human counterparties is naturally metered per verified entity — a fee each time the packager stands up a fresh evidence bundle proving who's really on the other end of a payment — with a subscription tier for institutions verifying at high volume who want a predictable bill instead of a variable one.
SMB cash-flow underwriting facts, built from consented bank data with adverse-action-ready reason codes, is where outcome-warranted pricing belongs: charge per underwriting decision the facts actually support, with a defined liability position if a reason code turns out to be wrong or unsupportable under ECOA scrutiny — the same logic that makes a mortgage QC warranty worth more than a subscription to an extraction API.
And the connective tissue between all three — the one-time integration fee for actually plugging into a bank's core system, a payment network's mandate API, or an EUDI Wallet relying-party integration — is usually the hardest six weeks of the entire sale. Pricing it at zero to win the logo doesn't save the customer money. It just moves that cost from their budget onto yours.
Treat everything in this section the way Chapter 1 insisted you treat any 2030 material: a plausible direction to position for, not a date to build against. Nothing below carries a citation, because nothing below has happened yet.
The building blocks are visibly assembling. A continuous, real-time audit — where a ledger is reconciled to source evidence constantly rather than sampled quarterly by a human auditor — becomes newly plausible once the underlying facts (bank transactions, invoices, payroll records) already carry the kind of provenance the EUDI Wallet and India's Account Aggregator network are built to produce. An auditor who currently samples a few hundred transactions and extrapolates could instead consume a standing feed of attestations and only intervene on exceptions. That is a genuine step change in how assurance work gets done, and it points to a live 2030 opportunity: sell the attestation layer, not the audit itself.
The more ambitious version is a portable "financial passport" — a verifiable-credential bundle of income, assets, and obligations that a person or a small business refreshes by consent and reuses across a mortgage application, a rental screening, and an SMB credit line, instead of re-proving the same facts to three different institutions in three different formats. The pieces exist in isolated form today: EUDI Wallet credentials, India's Account Aggregator and DigiLocker rails, and the KYC/KYB evidence packagers this chapter has already described as a 2027 build target. What doesn't exist yet is the interoperability and institutional trust to let one of those bundles walk unmodified from a bank's underwriting system into a landlord's screening tool into a different bank's SMB lending desk. That requires competing institutions to accept each other's evidence at face value, which is a trust problem, not a technology problem — and trust problems move on regulatory and reputational timelines, not sprint timelines. Anyone who tells you the financial passport ships broadly by a specific quarter is selling something. The honest plan is to build the evidence infrastructure now, in whichever single vertical you already touch, in a form that's portable in principle even before anyone else agrees to accept it.
Which is exactly the version of this story mortgage lending is already living through, one step ahead of the rest of finance. A wallet answers who gets to spend, on what authority, with what evidence behind the mandate. A mortgage answers a harder question sitting right underneath it: once the money is authorized, who gets to decide what it's actually for, and whose judgment does a lender trust enough to let the loan close without a human re-checking every fact one more time. That's Chapter 3.
For a decade, the mortgage question that made someone money was: what does this document say. The question worth paying for now is: given what these documents say, is this transaction acceptable, and what has to happen before it can close.
That is not a small shift in phrasing. It is the whole industry's moat moving to a different floor of the building. Extraction — turning a scanned paystub or a bank statement into structured fields a computer can use — is becoming a utility, priced like one and improving on someone else's roadmap, not yours. The layer above it is a different story: reconciling a borrower's paystub against their W-2 against their bank deposits against the county's recording requirements against the investor's overlay against the lender's own risk appetite, and producing a verdict that a regulator will accept eighteen months later if the loan goes bad. That layer has a name now. Call it the Decision Graph — a versioned, executable map of which rules apply to which loan, at which point in time, with which exceptions, and what happens when one of them fails. It is the thing an origination agent and an underwriting agent actually negotiate over when they sit down — machine to machine — to get a loan approved. It is also the thing that survives being commoditized, because it was never made of documents in the first place.
Look at where the price of "reading a mortgage document" actually sits in September 2026. Reducto, one of the sharper document-parsing platforms built for LLM pipelines, lists its standard extraction rate at $0.015 per credit — one credit per page per operation — with volume discounts below that for larger contracts 20. IBM shipped Docling for watsonx as a fully managed service in June 2026, built on an open-source toolkit that had already passed seventy million downloads, priced at roughly four dollars per thousand pages 21. Hyperscaler OCR sits underneath both of them as the free-to-cheap floor. This isn't one vendor losing a price war — it's an entire capability, turning a PDF into JSON, falling to commodity pricing inside about eighteen months, the same pattern the opening chapter of this book describes happening to one wave after another: the plumbing compounds quietly and gets cheap fast, while the exciting front-end bet is what everyone keeps arguing about.
Even the market-sizing exercises around intelligent document processing can't agree on how big the category is, and the disagreement itself is the signal. One widely cited estimate puts the IDP market at $4.3 billion in 2026, growing to $32.9 billion by 2033. Another puts it at $1.45 billion in 2026, growing to just $2.02 billion by 2033 22. That is roughly a 3x disagreement on the starting number and a completely different growth story layered on top of it — one analyst sees a category about to be the next infrastructure layer of finance, another sees a mature tool market growing in line with inflation. Neither estimate is obviously fraudulent; they are almost certainly counting different things under the same label. The lesson for anyone building here is not "trust the bigger number" or "trust the more conservative one." It's that a market-sizing slide should never be load-bearing in your plan. If two respected research houses can't agree within a factor of three on how large the pond is, you should be building for a specific defensible catch, not for the pond's abstract size.
That specific catch is visible in who is actually raising money in this space right now, and what they're selling. Sela raised $21 million across seed and Series A rounds in 2026 for voice agents that run mortgage sales calls end-to-end, handling borrower conversations and escalating to a human loan officer only when needed; the company says six of the ten largest independent mortgage banks in the U.S. now run it in production, with a $10 million annualized run rate reached in eighteen months 23. Balerion AI raised a $6 million seed to build a "unified reasoning engine" that analyzes an entire loan file holistically rather than field by field, targeting the roughly $12,000 average cost to originate a loan that manual processes and fragmented systems currently produce 24. Friday Harbor built an AI pre-underwriting platform and, in September 2026, integrated it directly with Freddie Mac's own Income Calculator API, so a lender can get a calculated qualifying income back in minutes with representation-and-warranty relief attached, before the file ever reaches a human underwriter 25. Every one of these companies is selling past extraction. None of them is pitching "we read your documents better." They're pitching a verdict, or the automation of the conversation that leads to one.
None of that means extraction skill is worthless — someone still has to feed clean, tamper-checked, cross-reconciled facts into whatever makes the decision. It means extraction alone stopped being a business model roughly the moment a hyperscaler could match your accuracy for a fraction of a cent per page. The business that survives sits one layer up, in the graph of rules that decides what those facts mean.
Picture the actual exchange, stripped of the marketing language both sides will eventually put around it. An origination agent — working on behalf of a broker, or increasingly a borrower's own assistant — has assembled a loan package: income documents, asset statements, the property address, the requested product. It doesn't want a document review. It wants an answer to one compound question: given this borrower, this property, this jurisdiction, this investor, and this lender's own overlays, is this transaction acceptable, and if not, exactly what needs to change to make it acceptable.
It sends that package, tagged with the context that actually determines the answer — state, county, transaction type, investor, lender — to an underwriting agent sitting on top of a lender's or aggregator's Decision Graph. What comes back is not a dump of the applicable regulations. It's a decision: approved, denied, or review required, attached to specific findings ("escrow amount inconsistent with the applicable requirement, high severity"), a short list of missing items or required remediation ("recalculate escrow, regenerate the closing disclosure"), and a set of evidence identifiers and rule versions the finding is traceable to. The origination agent doesn't get the underlying rule text. It gets the verdict, the reason coded to a rule ID, and the version of that rule that was in force on the date of evaluation.
That last withholding is the whole design decision, and what's actually being protected is not the regulation. California's escrow requirements are public. The CFPB's disclosure rules are public. Fannie Mae's Selling Guide is public. Anyone — including a well-resourced competitor with a scraper — can read the primary source material for free. What is not public, and what took years and thousands of adjudicated loan files to build correctly, is the interpretation: which of several overlapping rules actually governs a specific fact pattern, what takes precedence when a state requirement and an investor overlay disagree, which exceptions apply and under what conditions, how a rule's effective date interacts with a loan's lock date, and which validation actually catches the failure before it becomes a repurchase demand eighteen months later. That mapping — rule to jurisdiction to transaction type to document to field to condition to exception to validation to remediation — is the Decision Graph, and it is considerably harder to scrape than a knowledge base, because scraping only gets you the inputs. It doesn't get you the thousands of resolved edge cases that taught the system which input wins when two rules point in different directions.
This is also why exposing "search my regulations" as the product would be a mistake regardless of who builds it. A raw query interface hands a competitor your whole corpus one request at a time. A verdict interface — approved, denied, review-required, plus reason and remediation, plus a rule version and an evidence trail — gives a counterparty everything it actually needs to act, and gives away nothing it could use to rebuild your rule graph from the outside. That is not a technical nicety. It is the entire difference between a defensible business and a very well-organized free knowledge base with a login page.
None of this works if the request is routed the way a naive "MCP server full of mortgage knowledge" would route it — straight from an agent into a shared corpus. The layer that has to sit in between is what turns a lookup into a verdict:
The model engine's job in this picture is meaning, not measurement. The parts of the pipeline that touch escrow math, tolerance thresholds, and eligibility cutoffs stay deterministic on purpose, because "the LLM said it looked compliant" is not a sentence that survives a repurchase demand. The Decision Graph lives underneath the Knowledge Service and Rules Engine boxes — it's the thing that tells the Context + Policy Router which jurisdiction, investor, and lender rules even apply before either engine runs, and in what order they take precedence when two of them disagree.
A lender that answers "is this loan acceptable" is selling a verdict. A lender that answers "here is everything I know about escrow law" is giving away the only asset it has left.
Chapter 1 named the three things that create a moat once a protocol itself is free: verified facts with provenance, warranted outcomes, and accountable judgment. Mortgage lending is the cleanest place in this book to see all three attach to one artifact.
Verified facts with provenance means every value the Decision Graph produces carries its receipt: which document it came from, which page, which field, computed under which rule version, at what confidence, timestamped. An income figure isn't just a number — it's a number with a path back to the paystub, the W-2, the bank deposit that corroborates it, and the specific rule that governed how those three sources were reconciled into a single qualifying-income figure. That's expensive to build correctly and unglamorous to describe in a pitch deck, which is exactly the combination the opening chapter flags as defensible: nobody builds this for a demo, because a demo doesn't need to survive an audit.
Warranted outcomes means someone prices the cost of being wrong, not just the cost of answering. This is where mortgage lending offers the cleanest 2027-scale business inside the entire IDP category: post-close quality-control and repurchase-risk scoring, priced per loan with an actual defect warranty attached. Instead of selling "we'll tell you if the file looks compliant," the vendor sells "we'll tell you the probability this specific loan gets kicked back by the investor after purchase, and if we're wrong beyond an agreed threshold, we eat some of the cost." That's a fundamentally different sale than a subscription to an extraction API, and it's only possible once you have a Decision Graph precise enough to price its own error rate.
Accountable judgment is the piece people assume can't exist yet for a decision system, as opposed to a licensed human underwriter — and mortgage lending already has a working counterexample. Fannie Mae's Income Calculator gives lenders representation-and-warranty enforcement relief for the accuracy of an income calculation, on an income-source basis, when the lender uses the calculator's output as submitted 26. Freddie Mac's Asset and Income Modeler, built into Loan Product Advisor, does the same for income, asset, and increasingly employment verification — and Freddie Mac reports that loans originated using AIM are roughly 2.1 times less likely to produce a defect 27. Freddie Mac has also published that lenders who maximize this kind of automation originate loans that run about $1,500, or 14 percent, cheaper, with production cycles roughly five days shorter 28. What that actually means is this: two of the most conservative institutions in American finance have already agreed, in writing, in their own selling guides, to hold themselves accountable for a decision a system made, provided the system's inputs and rule version are documented well enough to audit later. This isn't hypothetical. It's a precedent already operating at scale, for the exact class of decision this chapter is about. The remaining question isn't whether accountable judgment can attach to a system's output — it's how far past income calculation that relief extends, and which vendor's Decision Graph earns it next.
A Decision Graph is not one product with one price. It is a ladder, and Chapter 0's menu maps onto every rung of it cleanly enough that the pricing model should be obvious the moment you know which rung you're selling:
| Tier | What the customer gets | How it's priced |
|---|---|---|
| Extract | Structured data from mortgage documents | Metered, per page or per document |
| Validate | Document and data validation against a rule set | Metered, per document, higher rate than raw extraction |
| Compliance | Regulatory and disclosure checks (TRID tolerance, escrow rules) | Subscription tier, rate-limited by monthly check volume |
| Eligibility | Investor and lender guideline matching | Subscription tier, one rung up, more rules exposed |
| Decision | Full transaction-level verdict with evidence and remediation | Per loan file, flat fee |
| Continuous | Re-checks automatically when a regulation or guideline changes | Subscription add-on to the Decision tier |
| Enterprise | Customer-specific overlays computed into the executable policy | Annual contract, one-time integration fee plus per-loan usage |
Moving down that ladder, the pricing model shifts from metered (Extract, Validate — commodity operations, priced by volume) to subscription (Compliance, Eligibility — ongoing access to a maintained rule set) to a per-transaction flat fee (Decision — a discrete verdict on a discrete loan) to a blended enterprise contract (a fixed integration fee for wiring in a lender's specific overlays, plus ongoing per-loan usage on top). Sell the wrong pricing model at the wrong rung — metering a Decision, or trying to flat-fee raw Extraction — and you either leave money on the table or price yourself out of a market that has already commoditized that rung.
The same ladder logic applies to who gets access, not just what they're charged. A Decision Graph exposed through role manifests — a loan officer's seat with roughly eight tools and a modest daily call limit, a processor's seat with around fifteen and a higher limit, an underwriter's seat with twenty-plus tools and the widest data access, a QC auditor's seat scoped narrowly to post-close review — is simultaneously a usability decision (nobody wants forty tools loaded into one agent's context) and a monetization decision (a lender pays more per seat for a role that touches more of the Decision Graph, the same way a database vendor charges more for an admin seat than a read-only one).
And sitting above every rung of the ladder is the one tier this chapter has already named as the real 2027 prize: post-close QC and repurchase-risk scoring, priced per loan with an actual defect warranty — outcome-warranted pricing, sold as an add-on to the Decision tier rather than a replacement for it, because a lender who already trusts your verdict enough to buy it is the easiest possible customer to sell a warranty on that same verdict to next.
The staging is what makes this a plan instead of a wish:
Extrapolate that precedent forward and the plausible 2030 picture comes into focus — everything from here is a bet on direction, not a reported fact. Picture a borrower's own agent holding a standing, continuously re-evaluated conditional approval rather than a one-time pre-qualification letter that goes stale the moment a pay stub changes. New income data lands, a Decision Graph re-runs the eligibility check automatically, and the approval updates itself — delivered not through a portal login but agent-to-agent, the same way a procurement agent checks a vendor's signed claims before opening a negotiation. Fannie and Freddie are already most of the way to the technical precondition for this: income and asset facts computed once, by an authoritative system, with relief attached to using them. What's missing for "always pre-approved" isn't the compute. It's the consented, refreshable data pipes and the regulatory comfort to treat a standing approval as something other than a one-time snapshot.
What is not plausible on any near-term timeline is the other half of that fantasy: an agent auto-locking a rate and closing a loan with no human ever touching the file. That's not a technology gap closing slowly — it's the Mandate Boundary sitting where the opening chapter said it would, drawn by the cost of being wrong rather than by what a model can technically execute. Preparing a decision, assembling evidence, and re-checking eligibility as facts change are all activities the boundary has already moved to accommodate, because a wrong pre-approval is embarrassing and correctable. Executing an irreversible six-figure financial commitment is not correctable in the same way, and TRID and ECOA disclosure requirements exist precisely to put a human signature between "the system recommends" and "the transaction is final." The pattern to build toward for 2030 is a lender's Decision Graph getting fast enough and trusted enough that the human signature becomes the very last, cheapest step in the process — not a pattern where that signature disappears. Watch for the boundary to keep moving on preparation for years before anyone credible proposes moving it on execution, and be skeptical of any vendor whose 2027 roadmap quietly assumes otherwise.
The underwriting agent that decides whether a loan closes and the payer agent that decides whether a claim gets paid are running the same logic wearing different regulatory costumes — a package of facts, a rule graph nobody outside the institution gets to read directly, and a verdict with a reason code attached. Chapter 4 walks into the second costume, where the agent on the other side of the negotiation isn't deciding who gets a house. It's deciding who gets care.
Somewhere right now, a piece of software representing a surgeon's office is assembling the case for a hip replacement, and a piece of software representing the insurer is testing that case against a coverage policy — and it is entirely possible that no human being on either side has opened the patient's chart yet. That is not a bug in the system. As of January 2026, in a meaningful slice of American healthcare, it is the system, and for once it is the system on purpose, by federal order, on a published clock.
Call the back-and-forth between those two systems the Adjudication Loop: the automated exchange in which a provider-side agent assembles a clinical case, a payer-side agent tests it against coverage criteria, evidence gets requested and resupplied, and the loop keeps iterating — sometimes closing in seconds — before a human on either side is pulled in. Every other domain in this book has a version of software negotiating with software over money. This is the one where the thing being negotiated is a person's access to care, and where getting the loop wrong is not a quarter of lost revenue. It is a denied hip replacement, a delayed chemotherapy start, a patient readmitted because the skilled-nursing stay got cut short by an algorithm that never saw the discharge note. The stakes change the entire calculus of where the Mandate Boundary can sit.
Every other chapter in this book describes a Mandate Boundary moving because a vendor got ambitious or a market got competitive. Prior authorization is different: the boundary is moving because the government wrote the date into a federal rule and started enforcing it.
CMS-0057-F, the Interoperability and Prior Authorization Final Rule, carries real enforcement weight, not vendor-roadmap ambition — it is binding on Medicare Advantage organizations, state Medicaid and CHIP fee-for-service and managed-care programs, and Qualified Health Plan issuers on the federally facilitated exchanges 29. Two of its deadlines already fell before this chapter was written. Since January 1, 2026, impacted payers must decide expedited requests within 72 hours and standard requests within 7 calendar days, and a denial can no longer arrive as a bare "no" — it has to come with a specific reason 29. On March 31, 2026, those same payers filed their first public prior-authorization metrics — approval and denial rates, appeal outcomes, average turnaround — and they will keep filing them every year from here 30. A second, larger deadline is still ahead: by January 1, 2027, impacted payers must expose four FHIR-based APIs, including a Prior Authorization API that lets a provider's system submit a request and track its status electronically instead of by fax and hold music 29.
Two qualifications matter, because they show the boundary moving on a schedule that is deliberately narrower than the hype around it. First, drug prior authorization is explicitly excluded from CMS-0057-F — a companion rule, CMS-0062-P, proposed in April 2026, would extend similar electronic requirements to drug PA, but as of this writing it is still a proposal with a closed comment period and no finalized date 31. Second, traditional Medicare and standalone Part D sit outside the rule entirely. This is not "AI is coming for all of healthcare administration by 2027." It is a specific, government-forced modernization of the medical-benefit prior-auth pipeline for specific payer types, with drug PA deliberately carved out for later. Build your 2027 plan against the rule that actually exists, not the version of it that sounds better in a pitch deck.
What CMS actually built, whether or not that was the intent, is a forcing function for exactly the kind of infrastructure this book has been arguing you need everywhere else voluntarily: evidence with a timestamp, decisions with a reason code, and outcomes reported in public where a regulator, a journalist, or a plaintiff's attorney can find them. The Mandate Boundary in prior authorization is being drawn by statute before most organizations have finished drawing it by choice.
The clearest evidence that the software-hires-software mechanic is already operating at scale sits inside a single company's numbers. Cohere Health processes more than 12 million prior-authorization requests a year for more than 600,000 providers, on the strength of a $90 million Series C led by Temasek that brought its total funding to $200 million 32. Its CEO, Siva Namasivayam, has said publicly that the AI renders a real-time decision in roughly 85% of those cases 32. A smaller competitor, Anterior, raised a $20 million Series A led by NEA at a $95 million valuation on the same thesis a year earlier 33. This is production infrastructure processing a double-digit-million volume of live clinical decisions, unaccompanied, most of the time — not a demo.
Read the sentence that follows Namasivayam's 85% figure, though, and you find the actual shape of the Mandate Boundary in this domain: "No claim is denied exclusively by AI" 32. The remaining 15% of decisions — and, more importantly, every decision that trends toward a denial rather than an approval — gets a nurse or a physician's sign-off before it becomes final. That is not a rounding error in the rollout. It is the whole architecture.
The asymmetry here is real: the Adjudication Loop is free to say yes by itself. It is not free to say no by itself. The same model that approves a clean case at 2 a.m. is perfectly capable of flagging one as non-compliant with criteria — the constraint isn't technical, it's a liability allocation, and it maps precisely onto the Mandate Boundary from Chapter 1: the boundary does not sit at "is the decision hard" or "is the model accurate enough." It sits at "which direction does the error cut." An erroneous approval wastes money and can usually be caught and clawed back later. An erroneous denial withholds care from a person who may not have the runway to wait for an appeal. Software gets the mandate to act alone in the direction where being wrong is expensive but recoverable. It does not get the mandate to act alone in the direction where being wrong might not be recoverable at all.
This is also, increasingly, written into law rather than left to vendor discretion. California's SB 1120 — the Physicians Make Decisions Act — requires that a medical-necessity determination be made by a licensed physician or a licensed health professional competent to evaluate the specific clinical question, and it explicitly bars an algorithm from denying, delaying, or modifying care based on medical necessity, in whole or in part 34. The FDA's own January 2026 update to its clinical-decision-support guidance moves in a similar, careful direction: it grants a CDS tool enforcement discretion to surface a single recommendation instead of a menu of options — but only when a single option is genuinely the only clinically appropriate one, and only when a clinician retains the ability to independently review the basis for it 35. Regulators are not saying the AI can't be trusted to reason well. They are saying the party who absorbs the consequence of a wrong denial has to be a licensed person who can be named, tested, and held to account — and that requirement is likely to outlive any improvement in model accuracy, because it isn't really about accuracy. It's about who a court can call as a witness.
An approval is a mandate the software can execute alone. A denial is a mandate that still needs a name attached to it — because "the model decided" is not an answer a regulator, a plaintiff's attorney, or a frightened patient will accept.
Apply the book's framework — verified facts, warranted outcomes, accountable judgment — to prior authorization, and each leg gets sharper teeth than it had in procurement, wallets, or mortgage lending, because the cost of getting any one of them wrong is measured in patient harm rather than balance-sheet risk.
Verified facts with provenance means a clinical evidence packet that maps specific facts from the chart — the imaging finding, the failed conservative-treatment history, the lab value — directly onto the specific line item in the payer's coverage policy it is meant to satisfy, with the policy version and its effective date attached. This is the harder half of the market to build in right now: the payer side is already crowded with well-funded platforms like Cohere Health and Anterior, but the provider side — packaging a specialty practice's own documentation into a submission a payer's AI can approve on the first pass instead of bouncing back for more information — is comparatively open, and it's priced naturally per approved authorization rather than per API call, because the entire point of the product is getting to "yes" faster, not generating more traffic.
Warranted outcomes are not a nice-to-have differentiator here; they are close to existential. In every other domain in this book, a bad AI decision produces an unhappy customer or a write-off. In prior authorization, a bad AI decision produces a lawsuit with a name on it. A class action against UnitedHealth Group, filed in late 2023, alleges that its nH Predict algorithm denied post-acute care with an extremely high error rate and little human review, with more than 80% of those denials reportedly reversed on appeal; a federal magistrate ordered broad discovery into UnitedHealthcare's AI-driven claims process in March 2026, and separately, a U.S. Senate investigation found the company's Medicare Advantage post-acute-care denial rate rose from 10.9% in 2020 to 22.7% in 2022 as automation scaled 36. Whatever the eventual legal outcome, the reputational shape of that story — "the AI denied grandma's rehab" — is now a known, replayable narrative that every payer and every vendor in this space has to build against. A warranty that puts real money behind an accuracy claim is the only credible answer to a story that has already been told once, publicly, in court filings — not a sales gimmick in this market.
Accountable judgment is where this domain diverges most sharply from every other chapter in this book. Elsewhere, "accountable judgment" can be a warranty fund, an insurance policy, a named institution absorbing risk on a balance sheet. In prior authorization, in a growing number of states, it is a specific licensed human being, by statute. SB 1120 does not ask for a reviewable audit trail as a best practice — it makes the licensed clinician the load-bearing legal requirement for a denial to be valid at all 34. That is the Mandate Boundary rendered as literally as it will appear anywhere in this book: not a policy an institution sets for itself, but a signature a named, licensed person is legally required to provide before the decision means anything.
Run prior authorization through Chapter 0's menu and the pricing follows the same asymmetry this chapter has already described between an approval and a denial.
A provider-side PA evidence packager — mapping a specialty practice's own documentation onto a payer's coverage criteria — is priced through outcome-warranted pricing: a fee per approved authorization, not per submission attempted, because the entire value proposition is getting to yes on the first pass instead of bouncing back and forth through the Adjudication Loop. A vendor charging per API call here is pricing the wrong thing; a practice doesn't care how many calls it took, it cares whether the patient got approved.
A medical-records chronology and facts API — page-cited chronologies for claims, life underwriting, or legal use — is a cleaner fit for metered or tiered subscription access, because the buyer here (an insurer's claims desk, a life-underwriting team, a law firm) is consuming a volume of records rather than betting on a single yes/no outcome; price it per record processed with a subscription tier for high-volume buyers, the same shape as the document-extraction pricing in the mortgage chapter.
The warranty layer behind either product deserves its own line item, priced closer to an insurance premium than a software fee, given what this chapter has already shown about the cost of getting it wrong: a percentage of the authorization value or a flat per-decision premium that scales with the litigation exposure of getting a denial wrong, sold explicitly as protection against the nH Predict-shaped headline risk every payer and vendor in this space now has to underwrite against.
And the 2030 health passport — patient-held, consented facts that arrive pre-verified at the next Adjudication Loop — has an unusual buyer for a subscription: not the patient, who shouldn't have to pay to be believed twice, but whichever side of the transaction saves the most when the loop resolves faster, which is most often the payer avoiding a lengthy appeal or the health system avoiding a denied, unpaid claim. Price it as a per-query fee paid by the party whose Adjudication Loop just got shorter, not as a subscription billed to the person the whole system is supposed to be serving.
The most quoted extrapolation about AI and medicine right now is some version of "instant, personalized drugs by 2030." Treat that claim as wrong the way it's usually stated, because the evidence people cite to support it doesn't actually support the consumer-scale version of the story.
The case everyone points to is a real one, and it is genuinely remarkable: a CHOP and Penn Medicine team designed a bespoke base-editing therapy for an infant, known publicly as KJ, born with a life-threatening urea-cycle disorder called CPS1 deficiency. The team went from diagnosis to a custom therapy in roughly six months, gave the first dose in February 2025, and published the results in the New England Journal of Medicine that May 37. In February 2026, the FDA followed up with a draft "Plausible Mechanism Framework" guidance, which would let sponsors seek approval for these individualized, n-of-1 genetic therapies for well-characterized, ultra-rare genetic mutations without running a traditional randomized trial that a single-patient population could never support 38. That is a real and significant loosening of the regulatory path for a real category of medicine.
None of that is "instant," and none of it is consumer-scale. Six months is extraordinary by the standards of drug development; it is not instant by the standards of a patient in front of you today. The infrastructure required — a specialized academic center, a bespoke molecular design process, a case sufficiently well-characterized that "the plausible mechanism" holds up to FDA scrutiny — describes a handful of centers treating a handful of patients with specific, well-mapped mutations, not a pharmacy counter dispensing a genome-matched pill to anyone who walks in. By 2030, the realistic bet is platform-based n-of-1 genetic medicines — base editors, antisense oligonucleotides — reaching more ultra-rare, well-characterized mutations, faster and more cheaply than they do today, still delivered in months at academic centers, still nowhere near consumer scale. That is a genuine advance. It is not the sci-fi version, and a roadmap built on the sci-fi version will misallocate a decade of capital.
The more useful 2030 bet is unglamorous, and it follows directly from the machinery this chapter has already described — so treat it explicitly as this book's own extrapolation rather than established fact: a portable "health passport," a patient-held bundle of consented, verifiable facts — diagnoses, medications, prior-authorization history, the specific evidence that satisfied a specific payer's criteria last time — that travels with the patient the way a credit history travels with a borrower. The plumbing already points this way: the Patient Access API that CMS is requiring by January 2027 exists to let a patient's own data follow them between payers 29; a health passport is what you get when a patient, rather than a payer, controls that data and hands relevant, pre-verified slices of it directly into the next Adjudication Loop. The payoff is not a faster diagnosis. It's a prior-authorization request that starts most of the way to "approved," because the evidence the payer's agent needs to see was already verified and attached the last time this patient needed similar care. That is a plausible, buildable extension of infrastructure that will exist by 2027. It is not guaranteed, and nothing in current CMS rulemaking commits to it — it is this book's best read of where the pattern goes next, offered as a bet worth positioning for, not a promise.
This is the last domain in this book, and it is worth closing on the plainest version of the Mandate Boundary you will find anywhere in it. Across a procurement agent shopping from another agent, a wallet that authorizes its own spending, a mortgage decision engine that recommends before a underwriter signs, and now a claim that argues with itself before either side's human sees it, the pattern never actually changes shape. Software gets faster, cheaper, and more capable of preparing the case — assembling the evidence, running the comparison, drafting the recommendation — and a specific, named, accountable party still has to sign it, and that party's identity, not the software's confidence score, is what a court, a regulator, or a frightened family will ask for when something goes wrong. What changes from domain to domain, and what this book has tried to make legible chapter by chapter, is only where that signature sits today and what it would take to move it. The organizations that win the next four years will not be the ones who guessed correctly whether the boundary moves in 2027 or 2030. They will be the ones who spent this window building the verified facts, the warranted outcomes, and the accountable judgment that let them move the Mandate Boundary deliberately, on evidence they can defend — instead of discovering where it has to sit only after a regulator, a competitor, or a courtroom moves it for them.
The Mandate Boundary, the Vendor Handshake, the Policy Wallet, the Decision Graph, and the Adjudication Loop are original framings developed for this book and carry no citation. All 2030 material is explicitly flagged in the text as the author's extrapolation, not reported fact.
Anuj Sadani builds AI systems and the teams that build them. He spent a decade at NVIDIA, has worked inside AI-first organizations, and has led multicultural engineering pods across Europe — sixteen-plus years, most of them on the AI era's unglamorous half: turning hype into systems that ship, and systems that ship into outcomes that hold up in production. That work has fed into industry recognition from Gartner and Everest Group, and an Innovator of the Year nod.
He is the author of The Clean Vibe Coder: A Code of Conduct for Programmers in the Age of AI Agents and Borrow the Line. Own the Move., and writes about engineering, leadership, and what stays human when the tools get good. He still believes the best technology is the kind that makes the people around it braver.
anujsadani.in
tech.anujsadani.in