Net Revenue Retention for Agents: Does Usage Expand or Collapse?
Net revenue retention is the single best leading indicator of whether an Agentic AI-as-a-Service business compounds or quietly bleeds out. But the SaaS version of the metric breaks the moment you bolt it onto usage-based, per-task agents. Expansion no longer comes from upsells and seat growth -- it comes from agents being trusted with more work, more often, across more workflows. Collapse rarely shows up as a cancellation; it shows up as a customer who quietly routes the hard tasks back to humans. This piece breaks down how to actually measure NRR for agents, why the number lies if you read it like a SaaS dashboard, and the specific mechanics that decide whether a cohort expands past 130% or rots below 80%.
Table of Contents
- Why NRR Is the Metric That Matters for GaaS
- The SaaS Definition Breaks for Agents
- What Actually Drives Agent Expansion
- The Three Faces of Agent Contraction
- How to Calculate Agent NRR Without Fooling Yourself
- Benchmarks: What Healthy Looks Like in 2026
- The Margin Trap Hiding Inside "Good" NRR
- Insights Most People Overlook
- References
Why NRR Is the Metric That Matters for GaaS
Net revenue retention measures what a cohort of existing customers spends this period versus what the same cohort spent a year ago, before you add a single new logo. It captures churn, contraction, and expansion in one number. A cohort that paid $1M last year and pays $1.3M this year -- without any new customers counted -- has 130% NRR. That is the holy grail of software, because it means the business grows even if sales stops selling.
For traditional software, NRR became the north star precisely because it predicts compounding. Bessemer's State of the Cloud research has spent years showing that the top public cloud companies cluster around or above 120% NRR, and that the metric correlates with durable growth far better than raw logo count. It's a quality-of-revenue signal. High NRR means customers are getting more value over time and paying for it.
For Agentic AI-as-a-Service, NRR matters even more, and for a slightly different reason. Most GaaS businesses don't sell seats. They sell completed tasks, resolved tickets, booked meetings, shipped pull requests -- units of work. (If that framing is new to you, the companion piece on cost-per-completed-task as the category's core unit lays out why the task, not the seat, is the atom of agent economics.) When revenue is tied to work volume, the question of whether a customer does more work through your agent next year is the entire ballgame. There's no contracted floor to fall back on. Expansion isn't a nice-to-have layered on top of a stable base. Expansion is the base, or there is no base.
That's why I'd argue NRR is the most honest GaaS metric we have. ARR is a fiction when revenue is lumpy and usage-based -- a point hammered home in the cluster piece on why per-task pricing makes forecasting nearly impossible. But a cohort either grew its spend or it didn't. NRR can't be dressed up the way an annualized run-rate can.
The SaaS Definition Breaks for Agents
Here's the problem. The textbook NRR formula was built for a world of fixed subscriptions and predictable upsell motions. Three things shatter when you apply it to agents.
First, the time window assumes spend is roughly continuous. SaaS customers pay every month whether or not they log in. Agent customers pay when they run tasks, and task volume can swing 4x between a quiet week and a launch week. Pick a bad measurement window -- say, you anchor your "starting cohort" during a customer's seasonal peak -- and you'll manufacture fake contraction in the comparison period. The volatility that makes token budgeting hard makes NRR windowing hard too.
Second, the SaaS formula treats a dollar of expansion as a dollar of expansion. But agent revenue carries wildly variable cost. A customer who triples their usage by sending you a flood of cheap, single-shot tasks expands your revenue and your margin. A customer who triples usage by handing you gnarly, multi-step jobs that fan out into dozens of sub-agent calls might expand revenue while destroying gross margin. Top-line NRR can't see the difference. You need a margin-aware view, which I'll come back to.
Third -- and this is the subtle one -- SaaS contraction is loud. A customer downgrades a plan or cancels, and it hits your billing system as an event. Agent contraction is silent. A customer doesn't cancel; they just send fewer hard tasks. The agent is still "live." The integration is still connected. The logo is still on your wall. But the revenue is quietly decaying because the human team stopped trusting the agent with the work that mattered. By the time it shows up in NRR, the relationship may already be unrecoverable. The cluster piece on why churn is invisible in GaaS until it's catastrophic is essentially about this failure mode.
So the SaaS definition isn't wrong, exactly. It's just blind to the three things -- volatility, cost variance, and silent contraction -- that decide whether an agent business lives or dies.
What Actually Drives Agent Expansion
If expansion is the base, you should know precisely where it comes from. In my read of how the better GaaS businesses actually grow, agent expansion has four distinct engines, and they're not the SaaS ones.
Trust-driven task migration. This is the big one. A customer starts by letting the agent handle the easy 30% of a workflow -- the password resets, the lead enrichment, the boilerplate code. As the agent proves reliable, the human team migrates harder tasks to it. Each migrated task category is real expansion. The whole motion is gated by reliability, which is why I think human-intervention rate is the leading indicator that predicts NRR a quarter or two ahead. When intervention rate falls, task migration accelerates, and expansion follows.
Autonomy expansion. Related but distinct: the same task, done with less human review. An agent that needed a human to approve every action becomes one that's trusted to run unsupervised on a class of work. Higher autonomy means more billable completions per unit of human oversight, and it often unlocks entirely new use cases the customer wouldn't have attempted with a human in the loop.
Workflow sprawl. The agent earns its way into adjacent workflows. A support agent that nails tier-1 tickets gets pointed at tier-2, then at internal IT requests, then at partner onboarding. This is the land-and-expand motion GaaS shares with SaaS, except the "expand" is measured in task types absorbed, not seats sold.
Volume growth in the customer's own business. If your agent handles a customer's outbound and their pipeline doubles, you bill more without doing anything. This is the most fragile engine, because it's not your win -- it disappears the moment the customer's market turns.
The first two engines are the durable ones, and notice they share a root cause: reliability. Agent NRR is, to a first approximation, a lagging measure of how much your customers trust your agent. a16z's writing on the emerging economics of AI agents makes a similar point from the pricing side -- outcome-based models only expand when buyers believe the outcomes are real and repeatable.
The Three Faces of Agent Contraction
Contraction is where the silent killers live. I see three patterns, in rough order of how dangerous they are.
Quiet task retreat
The customer stops sending the hard tasks. Volume on easy work holds, so total usage dips only slightly -- maybe 10% -- and it doesn't trip any alarm. But the mix has shifted toward low-value work, and the high-value tasks that justified the contract have gone back to humans. This is a leading indicator of churn dressed up as a healthy account. The only way to catch it is to track NRR by task category, not just in aggregate.
Margin contraction masquerading as expansion
Revenue grows, but the customer has discovered the agent is great at exactly the expensive, retry-heavy, fan-out-prone tasks. Your top-line NRR reads 115%; your gross-margin-weighted NRR reads 85%. You're "expanding" your way toward insolvency. The hidden cost of retries is the usual culprit here -- one task quietly becoming fifty model calls.
Hard churn
The integration gets ripped out. By the time this happens, one of the first two patterns almost always preceded it by a quarter or more. Hard churn in GaaS is rarely a surprise to anyone who was watching task-mix and intervention-rate trends. It's a surprise only to teams reading aggregate dollars.
The through-line: aggregate NRR is a trailing, lossy summary. The decision-grade signal lives one level down, in the composition of usage.
How to Calculate Agent NRR Without Fooling Yourself
Here's the approach I'd defend. Start from the standard formula and then fix the three things that break it.
The base calculation, applied to a cohort of customers who existed at the start of the period:
NRR = (Starting revenue + Expansion - Contraction - Churn) / Starting revenue
Now the fixes that make it honest for agents:
Use a trailing window, not a point-in-time snapshot. Define starting and ending "revenue" as trailing-3-month or trailing-6-month averages, not the spend in a single month. This neutralizes the usage volatility that would otherwise inject noise. If you compare December (a customer's peak) to a quiet June, you're measuring seasonality, not retention.
Compute a gross-margin-weighted NRR alongside the top-line one. Replace each revenue figure with its contribution margin (revenue minus the model, compute, and tool-call costs to serve it). When top-line NRR and margin-weighted NRR diverge, the gap is your early warning. A 30-point gap means your expansion is concentrated in your worst-margin work. (For how to even compute that per-task cost, the teardowns on a coding agent and a support agent at scale are the practical references.)
Decompose by task category, not just by customer. Track which categories of work are expanding and which are retreating. A customer can be at 100% aggregate NRR while their high-value task volume has fallen 40% -- masked by a surge in cheap work. The aggregate hides the rot. The decomposition exposes it.
One practical warning that the broader SaaS literature keeps relearning: be ruthless about how you handle one-time spikes. A consumption customer who ran a massive one-off backfill last year will look like catastrophic contraction this year when the backfill doesn't repeat. Strip non-recurring usage events out of the base, or you'll chase phantom churn. OpenView's work on usage-based pricing and its metric implications is the clearest treatment I've seen of why consumption revenue needs this kind of normalization before it means anything.
Benchmarks: What Healthy Looks Like in 2026
Numbers, with the caveat that public benchmarks for pure GaaS are still thin and self-reported. Treat these as directional, not gospel.
For best-in-class usage-based software broadly, the bar set by the cloud index is roughly 120%+ NRR for the leaders, with the median good performer in the 110-120% band. Agentic businesses that are working should run hotter than SaaS at this stage, not cooler -- because trust-driven task migration is a steeper expansion curve than seat expansion ever was. When an agent genuinely earns trust, a customer can 3x their task volume in a year. I'd want to see a healthy early-stage GaaS cohort north of 130% top-line NRR.
The number that should scare you is the floor. Because agent revenue has no contracted base, a GaaS cohort can post NRR below 80% far more easily than a SaaS cohort, where the subscription floor cushions the fall. Sub-100% NRR in GaaS isn't "slightly underperforming" -- it means your installed base is shrinking and you're on a treadmill, replacing decay with new logos. Given that outbound may not even work for selling agents, a treadmill is a death sentence.
The pairing that tells the real story: top-line NRR and margin-weighted NRR, side by side. 130% top-line with 125% margin-weighted is a fantastic business. 130% top-line with 90% margin-weighted is a company growing into a wall. The gap matters more than either number alone.
The Margin Trap Hiding Inside "Good" NRR
I want to dwell on this because it's the mistake I see ambitious GaaS teams make. They optimize for expansion. They celebrate the 130% top-line NRR in the board deck. And they don't notice that the expansion came from customers leaning into the agent's most expensive behaviors.
When you charge per task but your cost per task is variable -- swinging with retries, fan-out, thinking tokens, and tool-call stacking -- expansion can be margin-negative. A customer who doubles usage on a task type where your cost-to-serve is 90% of revenue has expanded your revenue and shrunk your gross profit dollars at the contribution level if it crowds out better work. NRR, the metric everyone trusts, cheerfully reports this as success.
The fix isn't to stop chasing expansion. It's to refuse to read top-line NRR in isolation. Every NRR review should put the margin-weighted version next to it, decomposed by task category, on a trailing window. That's three corrections to one metric -- and each one closes a specific way the SaaS definition lies about agents. Do all three, and NRR becomes the most trustworthy number on your dashboard. Do none, and it becomes the most dangerous.
Insights Most People Overlook
1. NRR is a trailing measure of trust, so your real leading indicator is intervention rate. By the time NRR moves, the customer's trust decision is months old. If you want to forecast next quarter's NRR, watch this quarter's human-intervention rate and task-migration velocity. Falling intervention rate predicts expansion before billing ever shows it. Most teams stare at the lagging dollar figure and get blindsided.
2. High NRR can be a margin liability, not an asset. The entire industry imported "NRR good, more NRR better" from SaaS without asking what the expansion costs to serve. In GaaS, the customers expanding fastest are often the ones who found your most expensive failure modes. Unqualified NRR worship will route your best engineering effort toward your worst-margin growth.
3. Silent contraction means your healthiest-looking accounts can be your most at-risk. An account at flat aggregate usage may have quietly migrated all its high-value work back to humans while keeping the cheap stuff on autopilot. Flat NRR can be the calm before the rip-out. The only defense is task-category decomposition -- aggregate dollars will lie to you right up until the cancellation email.
4. Usage volatility makes the measurement window a strategic choice, not an accounting detail. Pick the wrong anchor period and you can manufacture either fake expansion or fake contraction. Vendors who report NRR without disclosing their windowing methodology are, knowingly or not, giving you a number you can't trust. Always ask: trailing how many months, and how do you treat one-time spikes?
5. The GaaS NRR floor is structurally lower than SaaS, and that changes everything about defensibility. No contracted base means no cushion. A SaaS company at 95% NRR is mildly leaky; a GaaS company at 95% NRR is actively shrinking with nothing to catch it. This is why per-outcome and committed-use contracts are creeping back into "pure usage" GaaS pricing -- vendors are quietly rebuilding a floor because they've felt how fast the bottom falls out without one.
References
More in Economics
- Payback Period Math When Your Agent Revenue Is Usage-Based and Lumpy
- Why Per-Task Pricing Makes Forecasting Nearly Impossible (And What to Do Instead)
- The CAC Question: Does Outbound Even Work for Selling Agents?
- The Margin Trap of "We'll Just Pass Through Model Costs"
- Unit Economics Teardown: What a Sales-Development Agent Actually Costs at Scale