THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
Economics

Cohort Analysis for Agent Products: Why Retention Splits by Use Case

Most GaaS operators run a single retention curve across their whole customer base, and that curve lies to them. When you sell autonomous agents, retention is not a property of your product, it's a property of the *job* the agent does. A coding agent, a support-triage agent, and a research agent bought by the same logo will retain on completely different timelines, for completely different reasons. This piece shows how to build cohorts that respect use case, what curves you should actually expect, and the traps (survivorship, task-volume drift, "zombie autonomy") that make blended numbers dangerous.

By C. Whitlock · Jun 7, 2026 · 13 min read

Table of Contents

The blended curve is hiding three businesses

Here's a pattern I've watched repeat at agent startups that have crossed a few hundred paying accounts. Leadership pulls a monthly retention chart, sees something like 78% logo retention at month six, and files it under "fine." Then a quarter later revenue softens and nobody can explain it, because the blended curve averaged a use case retaining at 95% with one churning out at 40%, and the mix shifted.

This is not a new problem, SaaS people have warned about blended retention for two decades. What's new is that agents make the blending worse, because a single customer often runs your product across several jobs that have nothing in common. The same Series B fintech might use your agent platform to draft support replies, reconcile invoices, and write internal tooling. Roll those into one account-level retention number and you've created a statistic that describes none of the three.

The fix is conceptually simple and operationally annoying: stop cohorting by signup month alone, and start cohorting by what the agent was hired to do. The annoying part is that most teams never instrumented use case as a first-class dimension, so they can't reconstruct it. We'll come back to that.

Why use case is the right cohort axis for agents

Traditional cohort analysis slices by acquisition date, the January cohort, the February cohort, and tracks how each thins over time. That axis still matters. But for agent products it's the secondary axis. The primary one is use case, and here's the reasoning.

An agent's retention is governed by whether it reliably finishes the job it was bought for. That reliability is wildly job-dependent. A document-classification agent operating in a narrow domain might hit 97% task success on week one and never regress. A multi-step research agent that browses, synthesizes, and writes will have a success rate that wobbles with every model update and every change in the open web. Those two agents could be the same SKU, same pricing, same onboarding, and they will retain completely differently because the underlying task difficulty is different.

So when you cohort by use case, you're really cohorting by the thing that drives churn: how often the agent delivers a usable outcome versus how often it dumps work back on a human. That's tightly coupled to the human-intervention rate as a churn signal, when intervention climbs, the customer is quietly doing the job themselves, and the renewal is already lost even if the seat is still active.

There's also a topical-completeness reason to split this way. Retention by use case is the metric that connects unit economics to growth: it's where net revenue retention for agents gets decided, because expansion and collapse both happen inside a use case, not across your logo base. If you're building out a full GaaS metrics practice, this cohort table is the spine that the dashboard hangs off of.

What to actually count: tasks, not logins

The single biggest mistake in agent cohort analysis is importing the SaaS definition of "active." In seat-based SaaS, an active account is one whose users log in. For an agent product, logins are nearly meaningless, the entire value proposition is that the human doesn't have to show up. A perfectly healthy customer might never open your dashboard for weeks while the agent runs thousands of tasks on a schedule.

So your retention numerator has to be task-based. The cleanest definition I've seen in practice: a cohort member is retained in month N if it ran at least one successfully completed task in that month, in that use case, above some minimal volume floor. Two words in that sentence carry weight:

Andreessen Horowitz has argued that the right way to think about AI-native businesses is through the lens of work delivered rather than seats sold, and cohort retention is where that philosophy gets concrete. If your retention metric still counts heads, you're measuring the old business.

Building the cohort table step by step

The mechanics aren't exotic, it's the discipline around definitions that matters. Here's the build.

1. Define the cohort key

Each retained unit gets keyed on (use_case, acquisition_period). Use case has to be a stable, low-cardinality label, "support deflection," "code generation," "invoice reconciliation", not free-text. If a customer runs three use cases, they appear in three cohorts. Yes, that means one logo contributes to multiple curves. That's correct; you're measuring jobs, not logos.

2. Pick the period and the clock

Monthly periods work for most agent products. Start the clock at first successful task in the use case, not at contract signature. The gap between those two events is your time-to-value, and burying it inside an acquisition-date cohort hides activation failures that look like retention failures later.

3. Choose retention flavor

Run at least two curves per use case:

For usage-based agents these diverge hard. A cohort can show 90% logo retention and 130% revenue retention because survivors ramp task volume, or 90% logo retention and 70% revenue retention because survivors quietly throttled spend. The second case is the dangerous one, and a logo-only view is blind to it.

4. Normalize for task-volume drift

This is the agent-specific landmine. In SaaS, a seat is a seat. In GaaS, the amount of work per customer changes over time for reasons that have nothing to do with satisfaction, seasonality, a customer's own business growth, a workflow change upstream. If you don't normalize, you'll misread a customer whose business shrank as a churning customer, and vice versa. Track tasks-per-active-account inside each cohort as a companion series so you can tell "they stopped needing the work" apart from "they stopped trusting the agent."

Retention shapes you should expect by use case

Cohort curves have characteristic shapes, and after looking at enough agent products you start to recognize the archetypes. None of these are guarantees, they're priors to test against your own data.

The flat-and-sticky curve (narrow, high-reliability jobs)

Think classification, extraction, routing, bounded tasks where the agent either works or it doesn't, and when it works it keeps working. These cohorts drop in the first month or two as bad fits wash out, then flatten near-horizontal. Retention here is a qualification problem, not a durability problem: if you can get a customer past the first 60 days, they tend to stay for years. The strategic implication is to invest in onboarding and early-task success, because the back half of the curve takes care of itself.

The slow-bleed curve (open-ended, reliability-sensitive jobs)

Research agents, complex multi-tool workflows, anything where the task surface is unbounded. These curves don't fall off a cliff; they erode a few points every month as edge cases accumulate and trust slowly leaks. The killer here is that the erosion is invisible in any single month, it only shows up when you stack the cohort. This is the use case where churn is invisible until it's catastrophic, and the cohort table is your only early warning.

The expand-or-die curve (workflow-embedded jobs)

Coding agents and sales-development agents often look like this: a chunk of the cohort churns relatively fast (the agent never got embedded in the team's workflow) while the survivors expand aggressively (it became load-bearing infrastructure). The mean retention number is useless here because the distribution is bimodal. You have to look at the cohort as two sub-populations, and your whole growth model depends on widening the survivor share.

Gartner has noted that a large share of agentic AI projects will be scrapped before reaching durable production value, and that mortality concentrates in exactly these open-ended and embedded use cases. Your cohort table is where that industry-level attrition shows up at the account level, months before it shows up in revenue.

Reading the curve: expansion, collapse, and the silent middle

A cohort table is only useful if you act on what it tells you. Three signals to watch.

Expansion that's real versus expansion that's a billing artifact. Revenue retention above 100% feels great until you realize some of it is retry-driven, the agent is burning more tokens to do the same work, and you're billing for the inefficiency. That's not healthy expansion; it's the hidden cost of retries showing up as fake NRR. Cross-check revenue expansion against tasks-completed expansion. If dollars grow but completed tasks don't, you have a margin problem dressed up as a growth signal.

The silent middle. Most cohort attrition isn't a dramatic cancellation, it's a customer who drifts from 5,000 tasks a month to 500 over a quarter without ever filing a ticket. Account-level retention flags them as retained the whole way down. Only a volume-weighted, use-case cohort catches the slide while there's still time to intervene.

Curves that improve retroactively. When you ship a model upgrade or a reliability fix, watch whether older cohorts bend upward. They can, an agent that got more reliable can win back a customer who'd throttled usage. Blended numbers smear this out; use-case cohorts let you attribute a retention lift to a specific shipped improvement, which is gold for the product roadmap.

Instrumentation: the events you cannot reconstruct later

Everything above assumes you logged the right events. Most teams didn't, and you cannot backfill a use-case label onto a year of historical tasks. So, minimum viable instrumentation:

Stamp every task with a stable use_case identifier, a terminal outcome (completed / failed / escalated_to_human), and the customer/account it belongs to. Capture the escalation-to-human event specifically, it's the leading indicator that feeds your intervention rate and predicts churn one to two cohorts ahead of revenue. Log task volume per account per use case so you can normalize for drift. And timestamp first-successful-task per use case so your activation clock is honest.

If you do nothing else, capture outcome and use case on every task starting today. The cohort table you can build in six months is entirely determined by the events you start logging now. This instrumentation is also the raw feed for the GaaS metrics dashboard every operator should track, cohort retention isn't a standalone report, it's the most important panel on that board.

Insights Most People Overlook

References

#net revenue retention agents

More in Economics