THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
vs SaaS

Why Incumbents Have a Data Moat Agents Can't Easily Cross

The loudest version of the agent-disruption story says autonomous AI will hollow out SaaS by doing the work users used to do by hand. That's partly true. But it skips over an awkward fact: the incumbent that owns the system of record also owns the data the agent needs to be useful at all. An agent without proprietary, permissioned, continuously-updated context is a very expensive guesser. This piece breaks down why the data moat is real, where it's thinner than incumbents pretend, and how the smart agent vendors are crossing it anyway.

By A. Reyes · Jan 29, 2026 · 13 min read

Table of Contents

The Moat Everyone Forgets to Price In

Spend a week reading agent-disruption takes and you'll notice a pattern. The argument almost always runs through capability: the agent can read the email, draft the reply, update the CRM, file the expense, reconcile the ledger. Therefore the human doing those tasks is redundant, therefore the seats they occupied evaporate, therefore the SaaS vendor charging per seat is in trouble.

Every step in that chain assumes the agent can get at the data. That assumption is doing an enormous amount of unacknowledged work.

Consider a concrete case. You want an agent that handles your sales follow-ups end to end. To do that well it needs the full history of every deal, the notes your reps typed at 9pm after a bad call, the email threads, the contract redlines, the support tickets that flagged a churn risk, the entitlement data that says who's allowed to see what. Almost all of that lives inside Salesforce, or inside whatever system of record your company standardized on years ago. The agent's intelligence is a commodity you can rent from a foundation-model provider by the token. The data is not for rent. It belongs to the incumbent, and the incumbent decides the terms of access.

That's the moat. It's not a feature. It's a position.

What a Data Moat Actually Is (and Isn't)

The phrase "data moat" gets thrown around loosely, so it's worth being precise, because most of what people call a data moat isn't one.

Having a lot of data is not a moat. Public web text is enormous and protects nobody. Having data that's also available to your competitor is not a moat either. A moat requires data that is (a) proprietary, (b) hard to reconstruct from elsewhere, (c) continuously refreshed by ongoing use, and (d) wired into a permission structure the incumbent controls. The classic framing of data network effects, where each new user makes the product better for every other user through the data they generate, captures part of this, but for incumbents the more durable piece is simpler: they sit in the workflow, so the operational truth flows through them by default.

The distinction matters because it tells you which incumbents are actually defended and which just feel defended. A horizontal note-taking app holding generic meeting transcripts has a weak moat, that data is reconstructable and not deeply permissioned. A clinical records system holding decades of structured patient histories, tied to provider credentials and regulated access, has a deep one. Both have "lots of data." Only one has a moat an agent can't route around.

The Four Layers of an Incumbent's Data Advantage

Break the moat into its load-bearing parts and you can see exactly where it holds and where it cracks.

1. Proprietary Historical Data

This is the obvious layer and the most overrated in isolation. Years of transactions, records, configurations, and edge cases the incumbent has accumulated. It's valuable because an agent trained or grounded on it inherits institutional memory it could never derive from first principles. But raw history alone is a static asset. If that's all an incumbent has, a determined competitor can sometimes buy, scrape, or reconstruct an approximation. History matters most when it's combined with the layers below.

2. Real-Time Operational State

This is the underrated one. The incumbent doesn't just hold what happened, it holds what is happening right now. The current inventory count. The live order status. The open ticket queue. The in-flight approval. An agent that wants to act rather than just summarize needs the present-tense state of the business, and that state is being mutated continuously inside the incumbent's system. You cannot snapshot your way around this. The moment you copy the data out, it's stale. This is the heart of the system-of-record versus system-of-action fight, and it's why "just give the agent a data export" never produces a working autonomous workflow.

3. The Permission and Identity Graph

Who is allowed to see what, do what, approve what. Enterprises spend years encoding this, role hierarchies, field-level security, regional data residency rules, segregation-of-duties constraints. An agent operating in a real company has to respect all of it or it's a compliance incident waiting to happen. The incumbent already knows that this user can refund up to $500 but not $5,000, that this record is visible to the EU team but not the US team. Rebuilding that graph from scratch is brutal, and getting it wrong is dangerous. This layer quietly does more to lock agents out than the raw data ever does.

4. The Feedback Loop

Every correction a user makes, every "no, do it this way" inside the incumbent's product, becomes training signal the incumbent captures and the challenger doesn't see. This is the layer that compounds. An incumbent that ships its own agents on top of its own data gets a private flywheel: usage generates corrections, corrections improve the agent, a better agent drives more usage. As a16z has argued in its work on data and AI defensibility, the durable advantage in an AI product is rarely the model, it's the proprietary loop the product wraps around the model.

Why Agents Run Into the Wall

Put those four layers together and you can predict exactly where a standalone agent vendor hits resistance.

The agent can be brilliant and still be blind. It can have a frontier model's reasoning and zero idea that your top account just opened a critical support ticket, because that fact lives in a system it can't read in real time. It can draft a perfect contract clause and have no way to know your legal team banned that clause last quarter, because that decision was a correction logged inside the incumbent's tool.

There's also a deliberate dimension to the wall, which is its own battleground in this cluster: incumbents can simply gatekeep the integration. Rate-limit the API. Charge punitive fees for programmatic access. Change terms of service to forbid "competing AI agents." Bury the data behind a UI that has no clean machine interface. None of this requires the incumbent to out-innovate the agent. It just requires them to own the pipe and tighten the valve, the subject of the broader data-access wars playing out across the industry. Twitter's API repricing and LinkedIn's long war on scrapers are the pre-agent rehearsals for exactly this move.

So the wall has two bricks: data the agent genuinely can't get, and access the incumbent won't grant. Both are real. Neither is permanent.

Where the Moat Is Shallower Than It Looks

Here's where most incumbent-bull takes get lazy, and where the contrarian money is.

A data moat protects the data. It does not automatically protect the workflow built on top of it, and increasingly the workflow is where the value and the margin live. If an agent can get read-and-write access to the underlying records, even through the incumbent's own sanctioned API, it can host the workflow itself and relegate the incumbent to a dumb storage layer. The moat keeps the agent from owning the data. It does not keep the agent from owning the user relationship and the system of action.

Three things are actively eroding the moat:

Regulation. Open-banking mandates, data-portability rules under GDPR, and interoperability requirements are forcing incumbents in finance, healthcare, and beyond to expose data they'd rather hoard. The European Commission's framing in its data strategy treats data lock-in as a problem to be regulated away, not a competitive right. Every mandated API is a drawbridge lowered.

The integration layer maturing. Standards like the Model Context Protocol are turning "connect an agent to a data source" from a custom engineering project into a configuration step. As connection gets cheap, the incumbent's ability to hide behind integration friction shrinks.

Customers who own their own data refusing to be held hostage. The data inside Salesforce belongs, contractually and morally, to the customer. When a CIO decides the agent gets access, the incumbent's leverage is mostly bluff. The moat protects against outsiders reconstructing the data; it's much weaker against the data's actual owner choosing to pipe it to an agent.

The honest read: the moat is deep on proprietary real-time, permissioned, feedback-enriched state, and surprisingly shallow on static historical records the customer can simply export.

How Agent Vendors Are Crossing It Anyway

Smart agent companies have stopped trying to drain the moat and started building bridges. A few patterns are emerging.

Partner, don't pillage. Many vertical agent startups now ride on top of the incumbent's API as a sanctioned application rather than trying to replace the data layer. They concede the system of record and compete on the system of action. The incumbent keeps the storage rent; the agent takes the workflow value. It's an uneasy truce, and it's everywhere.

Go where the data is fragmented. The moat is deepest where one incumbent owns a clean, consolidated record. It's weakest where the relevant data is splattered across a dozen tools that don't talk to each other. Agents thrive in that mess precisely because stitching fragmented data together is exactly the chore humans hate and incumbents never solved. The agent's value is the integration, not the storage.

Generate proprietary data of your own. The most defensible agent vendors aren't renting the incumbent's moat, they're digging a parallel one. Every task an agent completes produces outcome data: what worked, what got corrected, what the user accepted. Run enough volume and you accumulate a behavioral dataset the incumbent doesn't have, because the incumbent only sees the record, not the reasoning trace of how the work got done. This is the agent-native answer to the moat, and McKinsey's analysis of where AI value actually accrues repeatedly lands on proprietary workflow data as the differentiator over model access.

Win the buyer before the incumbent reacts. Owning the agent layer means owning the relationship with the human who used to open the dashboard. Get there first and you can relegate the incumbent to plumbing even while you depend on its data.

What This Means for GaaS Economics

For anyone pricing or building Agentic-AI-as-a-Service, the data moat reshapes the unit economics in ways the pure capability story misses.

If your agent depends on an incumbent's data, that dependency is a cost line and a risk line. The incumbent can reprice API access and compress your margin overnight, the platform-risk problem that haunts every business built on someone else's foundation. Per-outcome pricing looks great until the outcome requires data you're paying a gatekeeper to access. Your gross margin is partly hostage to a competitor.

This is why the GaaS vendors with the most durable economics are the ones that either (a) operate in domains where the data is fragmented enough that integration is the product, or (b) generate enough proprietary outcome data that they've built their own moat alongside the agent. The vendors most exposed are the thin wrappers that bring nothing but a prompt and a borrowed model, sitting on data they neither own nor can defend.

The takeaway for the broader disruption thesis is more nuanced than either camp admits. Agents won't simply steamroll incumbents, because the data moat is real and in some verticals it's deep. But incumbents won't simply be saved by it either, because the moat protects the wrong thing, the storage, not the workflow, and the workflow is what customers will increasingly pay for. The fight isn't "agents versus SaaS." It's "who owns the layer where the work actually happens, given that the data lives over here and the action lives over there." That question gets settled vertical by vertical, and the data moat is the terrain, not the verdict.

Insights Most People Overlook

The moat protects the past tense, not the present tense. Static historical data is the layer everyone points to and the one most easily exported, scraped, or made portable by regulation. The genuinely defensible layer is real-time operational state, and almost nobody in the disruption debate distinguishes the two. An incumbent bragging about "20 years of data" while exposing live state through a generous API has the moat exactly backwards.

Permission graphs lock out agents more effectively than data scarcity does. You rarely hear this discussed, but the role-and-entitlement structure an enterprise spent a decade encoding is a harder barrier for an autonomous agent than the data volume. An agent that can't safely respect field-level security can't be deployed at all, regardless of how much data it can technically see.

The incumbent's own agents may leak the moat faster than competitors could. When an incumbent ships agents on its own data, it teaches the market that agent-shaped consumption of that data is normal and safe. It also has to expose more programmatic, agent-friendly interfaces to make its own agents work, and those same interfaces lower the wall for everyone. Defending by agentifying can quietly commoditize the very access you were guarding.

Outcome data is the agent-native moat, and it's invisible to incumbents. The record system sees the result; the agent sees the entire reasoning trajectory, the dead ends, the corrections, the human overrides. That trajectory data is a fundamentally new asset class the incumbent's architecture wasn't built to capture. The agent vendors who hoard it are digging a moat the incumbent can't even see being dug.

"We own the data" is often a contractual bluff against the data's actual owner. Incumbents talk as if the data is theirs. Legally it's usually the customer's. The moat is strong against third parties and weak against a CIO who decides the agent gets a connection. Vendors who confuse those two scenarios over-trust the moat.

References

More in vs SaaS