THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
Economics

Why Per-Task Pricing Makes Forecasting Nearly Impossible (And What to Do Instead)

Per-task and per-outcome pricing decouples revenue from the predictable anchors finance teams rely on, seats, contracts, renewal dates. Instead, your top line becomes a function of how often customers happen to run agents, how many retries each task triggers, and how much value the agent produces in a given week. That's three layers of variance stacked on top of each other, and most GaaS operators model only the first. This piece breaks down why the math fights you, where the variance actually hides, and how a handful of teams are forecasting anyway.

By E. Marchetti · Jun 10, 2026 · 13 min read

Table of Contents

The Forecasting Problem Nobody Warned You About

Here's a scene that's playing out in a lot of agent startups right now. The founder closes a strong quarter, walks into a board meeting, and gets asked the most basic question in business: "What's next quarter look like?" And they genuinely don't know. Not in the hand-wavy "the future is uncertain" sense, they don't know within a factor of two.

That's not incompetence. It's the structural consequence of a pricing model the whole category has rushed toward. When you sell agents per task or per outcome, cost-per-completed-task, per resolved ticket, per qualified lead, per merged pull request, you've traded the predictability of subscription revenue for the alignment of usage-based revenue. The alignment is real and it's a good reason to do it. But the predictability you gave up was load-bearing for the finance function, and almost nobody priced that trade honestly going in.

SaaS spent fifteen years building a forecasting apparatus around one assumption: revenue is a stock, not a flow. You have N customers each paying a fixed amount, you know your renewal dates, and next quarter's number is this quarter's number plus net new minus churn. Per-task pricing detonates that assumption. Revenue becomes a flow you sample, and the flow's rate is set by the customer's behavior, not your contract.

Why Seats Were Easy and Tasks Are Not

A seat is a commitment. When a company buys 200 seats of a SaaS tool, they're pre-paying for capacity they may or may not use, and crucially, they keep paying whether they use it or not. That last part is what makes forecasting tractable. The revenue exists independent of activity. Shelfware is a customer-success problem, but it's a forecasting gift, it means your MRR is decoupled from daily behavior.

A task is the opposite. It's pay-as-you-go in the most literal sense. No task, no revenue. And the number of tasks a customer runs in any given month isn't something they committed to, it's an emergent property of their workload, their headcount, their own seasonality, and a dozen things you can't see. You're not forecasting your business anymore. You're forecasting their business, refracted through your agent.

This is why the category still has no clean equivalent to MRR, the metric that anchored a generation of SaaS forecasts simply doesn't have a stable referent when the unit of value is a completed task rather than an occupied seat. a16z's analysis of how AI is upending the SaaS playbook makes the point bluntly: outcome-based pricing realigns vendor and customer incentives beautifully, but it also "introduces revenue that is inherently harder to predict," and the companies adopting it are essentially re-learning financial planning from scratch.

The seat made the customer's commitment your revenue floor. The task makes the customer's behavior your revenue, full stop.

The Three Layers of Variance Stacked on Top of Each Other

The reason per-task forecasting is nearly impossible rather than just hard is that the uncertainty isn't one distribution, it's three, multiplied together. Most operators model the first, occasionally the second, and almost never the third.

Layer one: how many tasks will customers initiate? This is demand variance, and it's the layer everyone at least attempts. The trouble is that agent demand is far spikier than SaaS logins. A support agent's volume tracks ticket volume, which tracks product launches, outages, and Mondays. A coding agent's volume tracks sprint cadence. You're inheriting your customers' operational rhythm wholesale.

Layer two: how many model calls will each task actually consume? This is the layer that quietly destroys margin forecasts. A "task" is a billing abstraction; underneath it is a variable number of model calls, tool calls, and, critically, retries. One support resolution might take three model calls on a clean run and fifty when the agent gets confused, re-plans, and loops. If you price per completed task but pay per token, layer two is the difference between a 70% gross margin and an underwater one, and it varies within the same SKU depending on which tasks show up.

Layer three: what's the outcome rate? For per-outcome pricing, you only bill when the agent succeeds, a resolved ticket, a booked meeting. So your revenue is initiated-tasks × success-rate, and success rate itself drifts as customer inputs change, as you ship model updates, as edge cases accumulate. You've now made revenue a product of three independent-ish random variables. The variance of a product of random variables is wider than any one of them, which is the mathematical reason a forecast that looks reasonable on each input can be wildly off on the output.

Stack those and you get the core thesis: it isn't that any single number is unknowable. It's that the composition compounds the error bars until next quarter's revenue is a wide, fat-tailed distribution rather than a point estimate with a tidy plus-or-minus.

The Demand Problem: Usage Is Lumpy and Seasonal

Let's sit with layer one, because it's the one finance teams underestimate even when they think they've accounted for it.

Usage-based revenue is lumpy in a way subscription revenue never is. A customer who runs 10,000 tasks in March might run 2,000 in April because their own quarter ended, a key champion went on leave, or they paused a workflow during a migration. None of that is churn. The account is perfectly healthy. But your revenue from it dropped 80% month over month, and if you're reading the top line without the usage context, it looks like the floor is falling out.

This is also why churn hides until it's catastrophic in usage-based models. A customer doesn't cancel, they just quietly run fewer tasks, then fewer, and by the time the revenue decline is obvious in aggregate, the behavioral disengagement happened months earlier. The forecasting signal you actually need lives in usage cohorts, not in the logo count.

Bessemer's State of the Cloud research on usage-based pricing documented this years before agents arrived: consumption-revenue companies show meaningfully higher revenue volatility quarter to quarter, and the public ones trade through wider guidance ranges precisely because they can't pin the number down. The agent wrinkle is that agents add layers two and three on top, so GaaS volatility is consumption-SaaS volatility squared.

The seasonal layer is sneakier still. Your revenue inherits the seasonality of every customer's industry simultaneously. A retail-heavy support-agent book peaks in Q4 and craters in January. A tax-software customer's agent volume is a cliff every April 16th. Blend enough industries and some of it averages out, but early-stage GaaS companies are rarely diversified enough for the law of large numbers to rescue them. Three big customers in the same vertical means your forecast is really one vertical's forecast wearing a trench coat.

The Cost Problem: Your COGS Moves Too

Here's the cruel part: even if you nailed the revenue forecast, the margin forecast would still be a mess, because the cost side moves independently.

In SaaS, COGS was hosting, a slow-moving, largely fixed line you could forecast to the dollar. In GaaS, COGS is inference, and inference cost per task swings for reasons that have nothing to do with your pricing. Token prices change. Your model mix changes when you route harder tasks to a more expensive model. Retry rates change with input quality. The same nominal "task" can cost you wildly different amounts week to week.

And don't assume falling token prices save you, there's a well-documented reason cheaper tokens didn't lower agent bills: as per-token costs drop, agents are designed to think more per task (longer reasoning chains, more tool calls, bigger context), so cost-per-task stays stubbornly flat or rises even as the unit input gets cheaper. Your COGS forecast built on "tokens are getting 40% cheaper this year" can be completely correct on the input and completely wrong on the output.

So now both sides of the gross-margin equation are stochastic. Revenue is initiated × success × price; cost is initiated × (calls-per-task × cost-per-call). The two share the "initiated" term, which helps a little, but the success rate and the calls-per-task can move in opposite directions, a model update that improves success (good for revenue) might also increase reasoning per task (bad for cost), and your margin lands somewhere you didn't predict on either axis. This is the same dynamic that makes the token-volatility problem a budgeting nightmare in its own right.

How a Few Teams Forecast Anyway

None of this means you throw up your hands. The teams that forecast usage-based revenue credibly do a few specific things, and they're worth stealing.

They forecast at the cohort and usage level, not the account level. Instead of "we have 50 customers averaging $4k/month," they model "customers in usage-tier 3 run a median of X tasks with this decay curve." This is just cohort analysis applied to forecasting, it turns the lumpiness into a distribution you can actually reason about rather than a single average that's wrong every month.

They forecast a range, not a number, and they own the range publicly. The mature consumption businesses give guidance as a band and educate their boards on why the band is wide. A point estimate in a usage-based business is a lie with a decimal point. The discipline is to express the P10/P50/P90 and explain what moves you between them.

They build in usage commitments to create a floor. This is the quiet pricing move underneath a lot of "usage-based" agent deals: a minimum monthly spend or a pre-purchased credit pool. It re-introduces a slice of the seat-like predictability, a revenue floor you can forecast, while keeping the per-task alignment above the floor. McKinsey's work on pricing in the age of AI agents repeatedly lands on hybrid structures for exactly this reason: pure consumption is too volatile for either party's planning, and a committed-base-plus-usage model is where most durable deals settle.

They instrument the leading indicators. Initiated tasks lead billed tasks. Active workflows lead initiated tasks. Seats-with-an-agent-configured leads active workflows. The further upstream you can measure, the earlier your forecast updates, and the difference between a good GaaS finance team and a bad one is how many weeks of warning their dashboard gives them before the revenue actually moves.

What This Means for Pricing Design

The deepest lesson is that pricing and forecastability are the same decision, and most teams treat them as separate.

Every step you take toward pure per-outcome pricing buys you incentive alignment and costs you predictability. Every step back toward commitments and platform fees buys predictability and costs you alignment. There's no free lunch, there's a dial, and where you set it should be a deliberate choice informed by how much variance your business can actually absorb. A venture-funded company chasing growth can tolerate a wider band than a bootstrapped one that needs to make payroll on a forecast.

What you should not do is set the dial all the way to "pure per-task" because it sounds modern and customer-friendly, then act surprised when finance can't forecast and the board loses confidence in your guidance. The volatility was a design choice. Own it, price a floor under it, instrument the leading indicators above it, and report in ranges. That's the realistic state of the art, not a model that makes the uncertainty disappear, but a discipline that makes it legible.

Insights Most People Overlook

Your forecast error is structurally fat-tailed, not symmetric. Most teams model variance as if revenue could miss high or low by equal amounts. It can't. The runaway-retry scenario, the customer who 5x's their usage after a successful pilot, the seasonal spike, these create a right tail on cost and a fat tail on revenue that a normal distribution badly understates. If you're sizing infrastructure or cash runway off a symmetric band, you're under-provisioned for the cases that actually hurt.

Diversification is your real forecasting strategy, and you can't buy it on demand. The single fastest way to tighten your revenue band is more uncorrelated customers across more verticals, the law of large numbers doing the work your model can't. But that's a function of go-to-market over years, not a forecasting technique you can deploy next quarter. The implication: early-stage GaaS companies are structurally un-forecastable and should stop pretending otherwise to their boards. Honesty about it is more credible than false precision.

The "task" you bill is a fiction your forecast inherits. Because a task is a billing abstraction over a variable bundle of model calls, two customers buying "the same" task are buying products with different cost structures. Your blended margin forecast hides this until your mix shifts. A new enterprise customer whose tasks are 3x more complex can tank blended margin while every per-customer number looks fine, a classic blended-rate illusion.

Success-rate improvements can hurt your forecast, not help it. Counterintuitive but real under per-outcome pricing: a model upgrade that raises success rate raises billed volume, which sounds great, but if it also raises reasoning-per-task, your cost rises faster, and you've improved the product while degrading the unit economics. Teams celebrate the success-rate win and miss the margin loss because the two numbers live on different dashboards.

Consumption pricing transfers forecasting risk to whichever party is worse at absorbing it, usually you. Finance teams on the buy side hate consumption pricing precisely because it makes their budgets unpredictable, and they push that risk back onto you via spending caps and committed minimums. The negotiation over the floor isn't a pricing detail; it's a fight over who eats the variance. Walk in knowing that's the actual subject.

References

#per-outcome pricing economics

More in Economics