The Near-Zero Marginal Cost Workforce: What Agent Labor Actually Costs
When the marginal cost of an extra unit of cognitive work falls toward the price of a few API tokens, the basic supply curve of labor stops behaving the way two centuries of economics assumed it would. Agentic AI-as-a-Service (GaaS) sells that shift as a product: workforce capacity you rent per task or per outcome instead of per head. But "near-zero" is doing a lot of quiet lifting in that pitch. The real marginal cost of an agent doing useful, reliable work is small but stubbornly nonzero, and the gap between the inference bill and the all-in cost is where the actual labor economics live. This piece breaks down how to think about agents as a labor input, why the supply curve goes nearly flat, and where the near-zero story quietly breaks.
Table of Contents
- Why "Marginal Cost" Is the Right Lens
- Where the Near-Zero Story Comes From
- The Hidden Marginal Costs Nobody Prices In
- A Labor Supply Curve That Goes Flat
- Per-Task and Per-Outcome Pricing as Labor Markets
- What Cheap Cognitive Labor Does to Wages and Demand
- The Jevons Trap: Why Cheaper Doesn't Mean Less
- Insights Most People Overlook
- References
Why "Marginal Cost" Is the Right Lens
Most coverage of AI and jobs argues about headcount: how many roles vanish, how many appear, who retrains. That's a fine debate, and other nodes in this cluster handle it directly. But it skips the mechanism. The reason agents threaten to reorganize white-collar work isn't that they're smart. It's that their marginal cost behaves nothing like a human's.
A human worker has a high and fairly fixed marginal cost. Hire one more analyst and you owe a salary, benefits, a desk, onboarding time, and management attention whether they're busy or idle. You can't summon a tenth of an analyst at 2 a.m. and dismiss them at 2:04. Labor, in the classical model, comes in lumpy, expensive, supply-constrained units.
An agent inverts almost all of that. The fixed costs are real but front-loaded onto the model provider and the platform. For the buyer, the cost of the next task is close to the cost of the tokens it burns. That's the whole game. When economists say a technology is general-purpose, they usually mean it lowers the cost of some input across many industries at once. Agentic AI lowers the cost of applied cognition itself, sold as a service, metered by the task. Marginal cost is the right lens because it's the variable that moved by orders of magnitude, and everything downstream, pricing, wages, firm size, follows from it.
Where the Near-Zero Story Comes From
The near-zero claim rests on a genuinely dramatic cost curve. The price of a given level of model capability has fallen steeply and repeatedly. Inference that cost dollars per million tokens a couple of years ago now costs a fraction of that for equal or better quality, and a16z's well-circulated work on the the cost of intelligence falling roughly 10x per year captured the slope that made the whole GaaS thesis credible. If the unit cost of a "thought" keeps collapsing, then renting cognition by the task starts to look like renting electricity.
There's a real structural truth underneath the hype. Software has always had near-zero marginal cost to copy; the novelty is software that performs open-ended knowledge work, so the near-zero copying property now extends to tasks that previously required a salaried human each time. Spin up one agent or ten thousand; the platform replicates the worker, not just the file. That parallelism, the ability to instantiate labor on demand and in bulk, is the part with no human equivalent.
So the headline isn't wrong. It's just incomplete. "Near-zero" is a statement about the inference line item. The labor-economics question is what the fully loaded cost of reliable agent work turns out to be once you stop looking only at that line.
The Hidden Marginal Costs Nobody Prices In
Here's where the clean story gets messy. The token bill is the visible cost. It's also frequently the smaller one. Several real marginal costs scale with each additional unit of agent work, and they don't trend to zero:
- Verification and review. An agent that's right 95% of the time still produces a wrong answer one task in twenty. Someone or something has to catch it. In high-stakes domains, the human-in-the-loop review time per task can dwarf the inference cost. Cheaper the agent gets, the more the checking becomes the binding constraint.
- Retries, loops, and orchestration. Autonomous workflows aren't single calls. A real agent task may chain dozens of model calls, tool invocations, and self-corrections. The advertised per-token price hides the multiplier. A "one task" can quietly be fifty calls.
- Error externalities. A bad human decision is contained by the human's limited throughput. A bad agent decision can replicate at machine speed across thousands of cases before anyone notices. The expected cost of rare catastrophic errors is a genuine marginal cost of deploying autonomy, and it rises with scale, not falls.
- Integration and context plumbing. Feeding the agent the right data, permissions, and tools is recurring engineering work. It amortizes, but slowly.
- Trust and liability overhead. Who's accountable when the agent is wrong? That cost lands somewhere, usually as insurance, contracts, or slower adoption.
Put bluntly: the marginal cost of inference approaches zero, but the marginal cost of trustworthy outcomes does not. The gap between those two numbers is the actual business of GaaS, and it's why agent reliability and agent security are their own cluster topics rather than footnotes. The vendors winning on per-outcome pricing are the ones who've driven that gap down, not the ones with the cheapest tokens.
A Labor Supply Curve That Goes Flat
Now the interesting part for an economist. Picture the supply curve for a category of cognitive labor, say, drafting a routine contract, reconciling an invoice, doing first-pass research. Historically that curve slopes upward: more output requires more workers, and at some point you bid up wages to get them. Supply is constrained by the number of trained humans willing to work at a given price.
Inject agents and you bolt a nearly horizontal segment onto that curve at a low price point. Below the agent's effective per-task cost, supply becomes almost perfectly elastic: you can buy essentially unlimited units of that task at roughly constant marginal cost. The classic constraint, the scarcity of human hours, stops binding for the slice of work agents can do reliably.
A flat supply curve does specific, predictable things. It caps the price of any task that sits clearly inside the agent's competence. It makes the quantity demanded extremely sensitive to demand shifts, because you're no longer rationing scarce labor. And it relocates scarcity. The bottleneck is no longer "can I hire someone" but "can I specify, verify, and trust the output." Scarcity moves from labor supply to judgment, taste, accountability, and the proprietary context an agent needs. Those become the expensive inputs, which is the seed of the "human premium" discussion elsewhere in this beat.
This is also why the displacement question is genuinely hard to call. When supply goes flat, the effect on humans depends entirely on whether total demand for the task expands enough to offset the collapse in price per unit. Which brings us to pricing models, and then to Jevons.
Per-Task and Per-Outcome Pricing as Labor Markets
GaaS pricing is the visible surface of this labor economics, and it's worth reading as a market design choice, not just a billing detail. McKinsey's framing of the economic potential of generative AI across knowledge work put hard numbers on how much activity is automatable; per-task and per-outcome pricing is how vendors monetize that automatable slice directly.
Three pricing regimes map onto three different labor-market structures:
- Per-token / per-call is closest to a raw commodity input market. You pay for compute, you bear the orchestration and verification risk. This is renting a tool, not labor.
- Per-task starts to look like a piece-rate labor market, the gig economy's pay-per-delivery model, pushed to its logical end. You pay for a completed unit of work, and the vendor absorbs the retries and plumbing. The vendor's margin is the spread between its true marginal cost and the per-task price.
- Per-outcome is the most labor-like of all: you pay only when the work produces the result you actually wanted, a booked meeting, a resolved ticket, a collected receivable. This is the closest a vendor can get to selling effort that's accountable for results, and it only works if the vendor has crushed the reliability gap, because they're now eating the cost of every failure.
The migration up that ladder, from per-token toward per-outcome, is the GaaS industry slowly converting "cheap inference" into "priced labor." Each rung shifts more of the hidden marginal cost from buyer to seller. The vendors who can afford to sell per-outcome are signaling that, for their narrow vertical, they've gotten the true all-in marginal cost low and predictable enough to guarantee it. That's a much stronger claim than "tokens are cheap."
What Cheap Cognitive Labor Does to Wages and Demand
If a near-zero-marginal-cost substitute for a task exists, basic price theory says the wage humans can command for that exact task gets capped near the agent's effective cost. You can't durably charge fifty dollars for something a buyer can get for fifty cents at acceptable quality. This is the mechanism behind the deflation-of-professional-services thesis that other articles in this beat examine.
But "wages get capped on the automatable task" is not the same as "wages fall for the worker." The two diverge in ways that matter:
- Task unbundling. Most jobs are bundles of tasks. Agents flatten the supply curve for some of them. A worker's pay reflects the residual, non-automated bundle, plus whatever leverage they gain by directing agents. The wage effect depends on which tasks were the high-value ones.
- Complementarity premiums. When a complement to labor gets cheap, the value of the scarce co-input often rises. If agents make execution cheap, judgment, relationships, and accountability, the things that decide which work to do and who's responsible when it's wrong, can command more. This is why the "agent boss" role and the question of who captures the productivity gains are live debates rather than settled ones.
- The level effect on demand. Cheaper output usually means more output bought. Whether that re-expands human employment depends on demand elasticity, the Jevons question below.
The honest summary: near-zero marginal cost reliably compresses the price of automatable tasks, and reliably raises the return to whatever stays scarce. Who wins among workers depends on which side of that line their skills fall on. That sorting, not aggregate job counts, is the part the headcount debate keeps under-weighting.
The Jevons Trap: Why Cheaper Doesn't Mean Less
The most common error in agent-labor analysis is assuming that cheaper cognitive work means less total cognitive work, and therefore fewer humans. The Jevons paradox, the 19th-century observation that more efficient coal use increased total coal consumption, is the standard caution, and it applies with unusual force here.
When the marginal cost of a task collapses, demand for it can explode in ways that were never economical before. Plenty of valuable knowledge work simply doesn't get done today because a human can't justify the hours: the contract nobody reviews, the data nobody analyzes, the customer nobody follows up with, the market nobody researches. Near-zero marginal cost doesn't just automate the existing pile of work. It pulls a vast reservoir of latent work above the line where it becomes worth doing.
So the labor question is genuinely indeterminate from first principles. If the demand curve for a category of cognitive work is elastic enough, near-zero marginal cost expands total work so much that human roles around the agents grow even as the per-task price craters. If demand is inelastic, fixed, capped, already saturated, then the same cost collapse mostly destroys roles. Different tasks will land on different sides of that line, which is exactly why blanket predictions about agent-driven employment tend to age badly.
The one safe generalization: near-zero marginal cost guarantees more cognitive work gets done. It guarantees nothing about who or what does the part that still requires a human. The economics of the GaaS workforce is, at bottom, a fight over that residual.
Insights Most People Overlook
-
The bottleneck migrates from supply to verification, and verification doesn't get cheaper. Everyone models the cost collapse on the production side. Almost no one models that when production goes to near-zero, checking becomes the dominant cost and the binding constraint. The most valuable human skill in an agent-saturated workflow isn't doing the task, it's cheaply and reliably knowing whether the agent did it right. Verification capacity, not labor supply, is the new scarce factor.
-
"Near-zero marginal cost" and "near-zero price" are not the same thing, and vendors profit from the confusion. The marginal cost of inference is collapsing. The price of reliable agent labor is set by what the reliability gap is worth, not by the token bill. A vendor with a 200x markup over inference cost can still be a bargain versus a human and a healthy business. The near-zero framing trains buyers to expect commodity pricing while the actual value, and margin, sits in the verification and trust layer.
-
Flat supply curves make demand volatility brutal. When labor was lumpy and expensive, demand swings were buffered by the cost of hiring and firing. When you can summon and dismiss ten thousand agent-workers instantly, there's no buffer. Capacity tracks demand in real time. That's efficient, but it also means the smoothing function that human labor markets quietly provided, the reason employment lags output, disappears. Expect cognitive-work "capacity" to become as spiky and financialized as cloud compute.
-
The first-order effect is deflationary, not inflationary, for affected services, and that's underpriced in macro forecasts. If a near-zero-marginal-cost substitute caps the price of a whole category of professional services, the effect on those prices is downward. Standard productivity-boom narratives emphasize growth; they under-discuss that agents may export persistent price deflation into the high-margin professional-services sector, with real consequences for an economy where those services are a large share of GDP.
-
Latent demand, not existing jobs, is the real prize, and incumbents are looking at the wrong number. Most analysis sizes the agent opportunity by counting current jobs and salaries that could be automated. That undercounts badly. The bigger market is the work that doesn't exist yet because it was never economical, the reviews, analyses, and follow-ups no human-hour budget could justify. The companies that win GaaS won't mostly be replacing today's labor line items; they'll be selling cognition for work that was previously priced out of existence.
References
More in Society
- The New Jobs the Agent Economy Is Quietly Creating
- The "Agent Boss" Is Already Here: What It Means to Manage a Fleet of AI Agents
- Which Roles AI Agents Augment vs. Replace: A Working Map for the Agent Economy
- What Agent Adoption Actually Does to Wages and Productivity
- The Job-Displacement Debate, Beyond the Hype