The Revenue-Quality Question: Is Usage Revenue Durable?
Agentic AI companies are booking eye-popping growth on usage- and outcome-based pricing, but investors are starting to ask the harder question: will those dollars still be there next year? Usage revenue can compound faster than any SaaS curve ever did, yet it can also evaporate the moment a model gets cheaper, a customer reroutes a workflow, or a pilot quietly dies. This piece breaks down what actually makes agent usage revenue durable versus disposable, the metrics that expose the difference, and why two companies with identical ARR can be worth wildly different multiples. The short answer: usage revenue is durable only when it sits on top of a workflow the customer can't easily reverse.
Table of Contents
- Why This Question Suddenly Matters
- Usage Revenue Is Not One Thing
- The Durability Spectrum
- Fragile Usage
- Sticky Usage
- The Metrics That Actually Reveal Durability
- The Model-Cost Deflation Trap
- How Investors Are Re-Underwriting Usage Revenue
- What Founders Can Do to Harden Revenue
- Insights Most People Overlook
- References
Why This Question Suddenly Matters
For most of the SaaS era, revenue quality was a solved problem. You sold seats, customers paid annually, and the renewal rate told you almost everything. A board could look at gross retention north of 90% and sleep fine. The contract was the moat.
Agentic AI broke that comfort. When you sell an agent that resolves support tickets, drafts contracts, or reconciles invoices, you're rarely selling seats. You're selling work performed, and you're often pricing per task or per successful outcome. That pricing model is genuinely better aligned with value, which is exactly why it spread so fast across the GaaS landscape. But it also means revenue is now coupled to volume of work, and volume of work is a far more volatile thing than a signed seat count.
The result is a class of companies posting growth numbers that look unreal. A vertical agent startup can go from a few thousand dollars a month to seven figures in a quarter, because once a customer trusts the agent on one workflow, the meter just runs. Investors love the curve. What keeps them up at night is whether the curve has a floor. If 40% of this quarter's usage is a Fortune 500 "exploring what's possible," that revenue has no floor at all.
This is the revenue-quality question, and it's now the central debate in how the market prices agent companies. It sits right next to the discussion of how the market assigns revenue multiples to agent companies and underwrites the entire funding conversation in this beat.
Usage Revenue Is Not One Thing
The biggest analytical mistake people make is treating "usage revenue" as a single category. It isn't. Two companies can both describe themselves as usage-based and have nothing in common from a durability standpoint.
Consider three flavors that frequently get lumped together:
Consumption you control. The customer has wired the agent into a production system, and usage scales with their business, not their curiosity. A coding agent that runs on every pull request, or a fraud agent that scores every transaction, generates usage that tracks the customer's underlying activity. This is the gold standard. It behaves almost like a subscription, except it grows automatically when the customer grows.
Consumption you trigger. The customer initiates usage deliberately, campaign by campaign or project by project. A marketing-content agent used during launches, or a research agent fired up for specific deals, produces lumpy revenue that can stay flat or vanish between cycles. It's real, but it's discretionary.
Consumption that's exploratory. Someone got a budget to "try AI agents," and the usage reflects experimentation rather than dependence. This is the revenue that looks identical to durable revenue on a dashboard and is the first to disappear in a budget review.
If you can't tell which bucket a dollar belongs to, you can't assess durability. And a lot of the ARR being celebrated in agent fundraising decks is a blend that's never been disaggregated, which is part of why agentwashing in pitch decks became a problem worth naming.
The Durability Spectrum
It helps to stop thinking binary and put usage revenue on a spectrum from fragile to sticky.
Fragile Usage
Fragile usage shares a few tells. It's concentrated in a handful of accounts. It came from a pilot or proof-of-concept that hasn't converted to a production commitment. It rides on a workflow the customer could perform another way tomorrow. And critically, it has no switching cost: if a competitor undercuts on price or a foundation-model lab ships a native feature that does 80% of the job, the customer can leave without ripping anything out.
A useful gut check: if the customer's CFO cut this line item, would anyone in the customer's org actually notice within a week? If the honest answer is no, that revenue is fragile no matter how much of it there is.
Sticky Usage
Sticky usage is the opposite, and it rarely comes from the agent's raw capability. It comes from everything that accretes around the agent over time. Sticky usage tends to involve accumulated state the customer would lose by leaving (a knowledge base the agent built, historical decisions, fine-tuned context). It's embedded in workflows that other systems and people now depend on. It often carries compliance or audit weight, where switching means re-validating an entire process. And it usually has multiple human stakeholders inside the customer who'd have to agree to remove it.
McKinsey's research on enterprise AI adoption keeps surfacing the same pattern: the value (and the lock-in) shows up not when a tool is deployed but when it gets rewired into core business processes, which is precisely where reversal becomes expensive. Durable agent revenue lives on the far side of that rewiring.
The Metrics That Actually Reveal Durability
Top-line ARR tells you almost nothing about durability. A handful of less glamorous numbers tell you nearly everything.
Net revenue retention, measured on usage. In SaaS, NRR above 120% is excellent. In durable usage businesses it can run far higher, because production usage expands naturally. But the number only means something if you can show it's driven by organic expansion within existing workflows rather than the same customers buying into new experiments. Andreessen Horowitz has written repeatedly that retention and expansion are the truest signal of whether AI revenue is real, and usage-based NRR is where that signal hides.
Cohort revenue curves, not blended. Blended growth masks churn. The honest view is each customer cohort's usage over time. Durable businesses show cohorts that flatten and then climb. Fragile businesses show cohorts that spike during onboarding and decay, the classic pilot-then-fade shape that no amount of new logos can fix.
Revenue concentration and the "production rate." What share of revenue comes from customers in genuine production versus pilots? A company that's 30% production and 70% pilot is a different risk than the inverse, even at identical ARR. This is the single most useful disaggregation, and it's the one founders are most reluctant to show.
Gross margin after inference. Usage revenue tied to model calls carries a variable cost that subscription revenue never did. If margins are thin because every dollar of revenue drags a heavy inference bill, the revenue is lower quality almost by definition, and it interacts dangerously with the burn-rate problem of running agents at scale.
The Model-Cost Deflation Trap
Here's the counterintuitive risk that catches even sophisticated investors. In most businesses, falling input costs are unambiguously good. In agent businesses priced per task on top of model APIs, falling model costs can shrink your revenue.
The mechanism is simple and brutal. If you price at a markup over inference, and inference prices drop 70% in a year, which has roughly been the trajectory for frontier-model token costs, then customers paying per task will expect that deflation passed through. Your unit revenue falls even as volume holds. You can lose top-line dollars while delivering more value than ever. This is why margin compression from cheaper models can force valuation haircuts on agent companies that look fine on a volume basis.
The companies insulated from this are the ones priced on outcome value rather than task cost. If you charge per resolved dispute or per closed candidate, the price floats with the customer's value, not your input cost. When the model gets cheaper, that's pure margin expansion, not revenue erosion. The pricing model you chose two years ago turns out to determine whether model deflation is your friend or your quiet killer. Anthropic's own guidance on building with usage and outcome pricing reflects this tension: the meter that's easiest to sell is often the one most exposed to deflation.
How Investors Are Re-Underwriting Usage Revenue
Late-stage diligence on agent companies now looks different than it did even a year ago. The smartest funds have stopped accepting ARR at face value and started demanding a usage-revenue teardown.
In practice that means asking for the production-versus-pilot split, requiring cohort curves rather than blended growth, stress-testing what happens to revenue if model prices fall another 50%, and probing concentration to see how much of the impressive number rests on two or three accounts that could each leave with a quarter's notice. This is the substance behind the broader shift in how VCs underwrite agent bets differently from SaaS, and it's a core item on any late-stage due-diligence checklist.
The valuation consequence is real. Two companies at $10M ARR can deserve a 3x spread in multiple based purely on revenue quality, one with durable production usage and outcome pricing, the other with fragile, deflation-exposed, pilot-heavy task revenue. The market is, slowly, learning to tell them apart. Bain's analysis of the generative AI market's commercial maturation points to exactly this kind of quality differentiation as the next phase of the cycle.
What Founders Can Do to Harden Revenue
If you're building a GaaS company, durability isn't something you discover at diligence, it's something you engineer from the pricing model up.
The highest-leverage move is pricing on outcomes rather than tasks wherever the value is measurable, because it decouples your revenue from model-cost deflation and aligns you with the customer's success. Beyond pricing, the durable companies deliberately build state that customers would lose by leaving: institutional memory, accumulated context, audit trails, and configuration that took months to tune. They push hard to convert pilots into production commitments with explicit volume floors or minimums, which transforms exploratory revenue into something with an actual floor. And they instrument their own usage data well enough to know, customer by customer, which bucket each dollar lives in, because if you can't see your own fragility, you certainly can't fix it before an investor finds it for you.
None of this is about being defensive. It's about recognizing that in agentic AI, the product is the easy part to copy and the embeddedness is the hard part to copy. The revenue that survives is the revenue that's been woven into something the customer can no longer easily unweave.
Insights Most People Overlook
Deflation can disguise itself as churn. When per-task revenue falls because model costs dropped and you passed savings through, your dashboards may show declining revenue per account and flag it as churn risk. It isn't churn, the customer is more committed than ever, but the metric lies. Teams that don't separate price-driven decline from volume-driven decline misdiagnose their healthiest accounts as their sickest.
The best usage revenue looks boring on the growth chart. Truly durable, production-embedded usage often grows steadily rather than explosively, because it tracks the customer's underlying business rather than their experimentation budget. The hockey-stick decks that excite investors are sometimes powered by exactly the fragile, exploratory usage that won't survive a budget cycle. Smooth, unsexy cohort curves can be the higher-quality asset.
Outcome pricing trades volatility for binary risk. Everyone praises outcome-based pricing for aligning incentives, but it has a hidden failure mode: if the agent's quality dips or the customer disputes attribution, revenue can drop to zero on workflows that were fully "live." Task pricing is volatile but rarely binary. Outcome pricing is durable when it works and catastrophic when the outcome definition gets contested, which makes agent reliability a direct revenue-quality input, not just an engineering concern.
A single foundation-model feature release can vaporize a revenue line overnight. Agent companies whose usage rides on capabilities the labs are likely to ship natively are holding deflation risk and obsolescence risk simultaneously. The durability question isn't only "will the customer stay", it's "will the thing they're paying me for still be something worth paying for after the next model release." Revenue sitting in the labs' obvious roadmap is fragile by construction.
Concentration hides inside "great" NRR. A company can post 140% net revenue retention while two accounts drive most of the expansion. The blended number looks like durability; the underlying distribution is a coin flip on two renewals. Always ask whether stellar retention is broad-based or a small number of whales, the answer changes the risk profile completely.
References
More in Market
- How VCs Are Underwriting GaaS Bets Differently From SaaS
- Seed-Stage GaaS: What Investors Actually Want to See Before They Write the Check
- Agentwashing: How to Spot the Fake "Agent" Startups Hiding in VC Pitch Decks
- Series A Benchmarks for Agent Companies: What It Actually Takes to Raise in 2026
- Why Agent Startups Command Premium Valuations (And When That Premium Is a Trap)