THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
Society

The Near-Zero Marginal Cost Workforce: What Agent Labor Actually Costs

When the marginal cost of an extra unit of cognitive work falls toward the price of a few API tokens, the basic supply curve of labor stops behaving the way two centuries of economics assumed it would. Agentic AI-as-a-Service (GaaS) sells that shift as a product: workforce capacity you rent per task or per outcome instead of per head. But "near-zero" is doing a lot of quiet lifting in that pitch. The real marginal cost of an agent doing useful, reliable work is small but stubbornly nonzero, and the gap between the inference bill and the all-in cost is where the actual labor economics live. This piece breaks down how to think about agents as a labor input, why the supply curve goes nearly flat, and where the near-zero story quietly breaks.

By M. Hale · Feb 4, 2026 · 13 min read

Table of Contents

Why "Marginal Cost" Is the Right Lens

Most coverage of AI and jobs argues about headcount: how many roles vanish, how many appear, who retrains. That's a fine debate, and other nodes in this cluster handle it directly. But it skips the mechanism. The reason agents threaten to reorganize white-collar work isn't that they're smart. It's that their marginal cost behaves nothing like a human's.

A human worker has a high and fairly fixed marginal cost. Hire one more analyst and you owe a salary, benefits, a desk, onboarding time, and management attention whether they're busy or idle. You can't summon a tenth of an analyst at 2 a.m. and dismiss them at 2:04. Labor, in the classical model, comes in lumpy, expensive, supply-constrained units.

An agent inverts almost all of that. The fixed costs are real but front-loaded onto the model provider and the platform. For the buyer, the cost of the next task is close to the cost of the tokens it burns. That's the whole game. When economists say a technology is general-purpose, they usually mean it lowers the cost of some input across many industries at once. Agentic AI lowers the cost of applied cognition itself, sold as a service, metered by the task. Marginal cost is the right lens because it's the variable that moved by orders of magnitude, and everything downstream, pricing, wages, firm size, follows from it.

Where the Near-Zero Story Comes From

The near-zero claim rests on a genuinely dramatic cost curve. The price of a given level of model capability has fallen steeply and repeatedly. Inference that cost dollars per million tokens a couple of years ago now costs a fraction of that for equal or better quality, and a16z's well-circulated work on the the cost of intelligence falling roughly 10x per year captured the slope that made the whole GaaS thesis credible. If the unit cost of a "thought" keeps collapsing, then renting cognition by the task starts to look like renting electricity.

There's a real structural truth underneath the hype. Software has always had near-zero marginal cost to copy; the novelty is software that performs open-ended knowledge work, so the near-zero copying property now extends to tasks that previously required a salaried human each time. Spin up one agent or ten thousand; the platform replicates the worker, not just the file. That parallelism, the ability to instantiate labor on demand and in bulk, is the part with no human equivalent.

So the headline isn't wrong. It's just incomplete. "Near-zero" is a statement about the inference line item. The labor-economics question is what the fully loaded cost of reliable agent work turns out to be once you stop looking only at that line.

The Hidden Marginal Costs Nobody Prices In

Here's where the clean story gets messy. The token bill is the visible cost. It's also frequently the smaller one. Several real marginal costs scale with each additional unit of agent work, and they don't trend to zero:

Put bluntly: the marginal cost of inference approaches zero, but the marginal cost of trustworthy outcomes does not. The gap between those two numbers is the actual business of GaaS, and it's why agent reliability and agent security are their own cluster topics rather than footnotes. The vendors winning on per-outcome pricing are the ones who've driven that gap down, not the ones with the cheapest tokens.

A Labor Supply Curve That Goes Flat

Now the interesting part for an economist. Picture the supply curve for a category of cognitive labor, say, drafting a routine contract, reconciling an invoice, doing first-pass research. Historically that curve slopes upward: more output requires more workers, and at some point you bid up wages to get them. Supply is constrained by the number of trained humans willing to work at a given price.

Inject agents and you bolt a nearly horizontal segment onto that curve at a low price point. Below the agent's effective per-task cost, supply becomes almost perfectly elastic: you can buy essentially unlimited units of that task at roughly constant marginal cost. The classic constraint, the scarcity of human hours, stops binding for the slice of work agents can do reliably.

A flat supply curve does specific, predictable things. It caps the price of any task that sits clearly inside the agent's competence. It makes the quantity demanded extremely sensitive to demand shifts, because you're no longer rationing scarce labor. And it relocates scarcity. The bottleneck is no longer "can I hire someone" but "can I specify, verify, and trust the output." Scarcity moves from labor supply to judgment, taste, accountability, and the proprietary context an agent needs. Those become the expensive inputs, which is the seed of the "human premium" discussion elsewhere in this beat.

This is also why the displacement question is genuinely hard to call. When supply goes flat, the effect on humans depends entirely on whether total demand for the task expands enough to offset the collapse in price per unit. Which brings us to pricing models, and then to Jevons.

Per-Task and Per-Outcome Pricing as Labor Markets

GaaS pricing is the visible surface of this labor economics, and it's worth reading as a market design choice, not just a billing detail. McKinsey's framing of the economic potential of generative AI across knowledge work put hard numbers on how much activity is automatable; per-task and per-outcome pricing is how vendors monetize that automatable slice directly.

Three pricing regimes map onto three different labor-market structures:

The migration up that ladder, from per-token toward per-outcome, is the GaaS industry slowly converting "cheap inference" into "priced labor." Each rung shifts more of the hidden marginal cost from buyer to seller. The vendors who can afford to sell per-outcome are signaling that, for their narrow vertical, they've gotten the true all-in marginal cost low and predictable enough to guarantee it. That's a much stronger claim than "tokens are cheap."

What Cheap Cognitive Labor Does to Wages and Demand

If a near-zero-marginal-cost substitute for a task exists, basic price theory says the wage humans can command for that exact task gets capped near the agent's effective cost. You can't durably charge fifty dollars for something a buyer can get for fifty cents at acceptable quality. This is the mechanism behind the deflation-of-professional-services thesis that other articles in this beat examine.

But "wages get capped on the automatable task" is not the same as "wages fall for the worker." The two diverge in ways that matter:

The honest summary: near-zero marginal cost reliably compresses the price of automatable tasks, and reliably raises the return to whatever stays scarce. Who wins among workers depends on which side of that line their skills fall on. That sorting, not aggregate job counts, is the part the headcount debate keeps under-weighting.

The Jevons Trap: Why Cheaper Doesn't Mean Less

The most common error in agent-labor analysis is assuming that cheaper cognitive work means less total cognitive work, and therefore fewer humans. The Jevons paradox, the 19th-century observation that more efficient coal use increased total coal consumption, is the standard caution, and it applies with unusual force here.

When the marginal cost of a task collapses, demand for it can explode in ways that were never economical before. Plenty of valuable knowledge work simply doesn't get done today because a human can't justify the hours: the contract nobody reviews, the data nobody analyzes, the customer nobody follows up with, the market nobody researches. Near-zero marginal cost doesn't just automate the existing pile of work. It pulls a vast reservoir of latent work above the line where it becomes worth doing.

So the labor question is genuinely indeterminate from first principles. If the demand curve for a category of cognitive work is elastic enough, near-zero marginal cost expands total work so much that human roles around the agents grow even as the per-task price craters. If demand is inelastic, fixed, capped, already saturated, then the same cost collapse mostly destroys roles. Different tasks will land on different sides of that line, which is exactly why blanket predictions about agent-driven employment tend to age badly.

The one safe generalization: near-zero marginal cost guarantees more cognitive work gets done. It guarantees nothing about who or what does the part that still requires a human. The economics of the GaaS workforce is, at bottom, a fight over that residual.

Insights Most People Overlook

References

#per-task agent pricing

More in Society