THE INDEPENDENT RECORD · AGENTIC AI AS A SERVICE AboutStandardsContact
GAASAGENTIC AI · AS A SERVICE
INDEPENDENT · SINCE 2026
UPDATED DAILY
NO HYPE · NO PAY-TO-PLAY
PER-TASK PRICING NOW STANDARD ● NEW BENCHMARK: 71% TASK COMPLETION ● ENTERPRISE PILOTS UP 4X ● RUNTIME FUNDING ACCELERATES ● "AGENTS ARE THE NEW SEATS" ● MARGINS UNDER PRESSURE ● THE INDEPENDENT RECORD ON GAAS
Market

The Late-Stage Investor's Agent Due-Diligence Checklist: What to Verify Before You Write the Check

Late-stage diligence on an Agentic AI-as-a-Service (GaaS) company looks nothing like a SaaS data-room exercise. The same dashboards that comfort a Series C lead, net revenue retention, gross margin, logo growth, can hide model-cost exposure, brittle autonomy, and "revenue" that is really one renewal away from churning. This checklist walks through what actually matters at the growth stage: revenue durability and quality, the real gross margin after inference, agent reliability and failure modes, security and liability surface, model-dependency risk, and the org's ability to keep an autonomous product alive in production. Treat it as a field guide for Series C and beyond, where valuations are high and the downside is a writedown.

By L. Karlsson · May 21, 2026 · 14 min read

Table of Contents

Why Late-Stage GaaS Diligence Is Different

By the time a GaaS company reaches Series C, the founder has a polished deck, a pipeline of recognizable logos, and a growth chart that bends in the right direction. The temptation is to run a standard SaaS playbook: confirm ARR, sanity-check NRR, model the path to profitability, and price the round off a revenue multiple. That playbook will get you in trouble here.

The problem is that an agent company sells outcomes, not seats. When you sell seats, revenue is a function of how many people log in. When you sell completed tasks, resolved support tickets, reconciled invoices, closed candidate screens, revenue is a function of how often the agent succeeds, how much each success costs to produce, and whether the customer could have done it another way once the novelty wears off. Each of those variables can move sharply and independently, and none of them shows up cleanly on a SaaS-style metrics dashboard.

Late-stage capital is also priced for a specific outcome: a large exit or IPO inside a few years, at a multiple that assumes durable, expanding revenue. If any of the agent-specific risk factors below turn out to be real, the company doesn't just grow slower, the unit economics can invert, and a growth-stage entry becomes a down-round candidate. The diligence below is built to find those failure modes before the term sheet, not after.

Revenue Quality: The First Real Test

The headline ARR number is the least interesting figure in the data room. What you want to know is whether that revenue is durable, whether it survives a renewal cycle, a budget review, and a competitor's pitch.

Start by decomposing the revenue by pricing model. Per-seat revenue behaves like SaaS and is the easiest to underwrite. Per-task and per-outcome revenue is where the value, and the risk, concentrates. Ask for cohort curves on consumption: do customers who signed 18 months ago run more tasks today, or have they plateaued? Consumption that flattens early is a tell that the agent solved a finite problem rather than embedding into a workflow that grows.

Then interrogate the "outcome" definition itself. Many GaaS companies book revenue on outcomes the customer cannot easily verify. If an agent claims to have "resolved" a support ticket, who decides it was resolved, the vendor's own classifier, or the customer's CSAT data? Revenue tied to a metric the seller controls is lower quality than revenue tied to a metric the buyer audits. Pull a sample of accounts and reconcile booked outcomes against the customer's own systems of record. Discrepancies here are a leading indicator of future disputes and churn.

A few quality flags worth scoring explicitly:

The deeper version of this analysis, modeling whether usage-based revenue holds up under a downturn, deserves its own workstream; treat the revenue-durability stress test as a mandatory companion to this checklist, not an afterthought.

The Margin Question Behind the Margin

Software investors are trained to expect 75-85% gross margins. Many GaaS companies report numbers in that range. The question is what's been excluded from cost of goods sold.

The single largest variable cost in an agent business is inference, the tokens consumed by the underlying model on every task, including retries, tool calls, and the multi-step reasoning chains that make agents "agentic" in the first place. A genuinely autonomous agent might call a frontier model a dozen times to complete one task. If those costs are parked below the line, or amortized in a way that flatters the current period, the reported gross margin is fiction.

Demand a COGS bridge that puts inference, vector-database queries, third-party tool/API fees, and human-in-the-loop review labor inside gross margin. Then look at the trend. A healthy GaaS company shows margins improving over time as it optimizes prompts, routes cheaper models to easier tasks, caches aggressively, and renegotiates model pricing. A dangerous one shows margins flat or compressing as task complexity grows, meaning it's buying revenue at a loss it hopes to engineer away later.

Two scenarios deserve explicit modeling. First, what happens to margin if frontier model prices fall (good for the company) versus if the company's task mix shifts toward harder, more token-hungry work (bad)? Industry analysts at firms like a16z have written extensively on how inference costs reshape AI application economics, and the punchline is that margin is a moving target, not a fixed input. Second, model the company's exposure to a single provider's price changes, which leads directly to the next section.

Reliability: Can the Agent Actually Be Trusted in Production?

This is the workstream most financial-only diligence teams skip, and it's often where the real risk lives. You need someone technical in the room asking how the agent behaves when things go wrong.

Ask for the production success rate, defined honestly: of all tasks the agent attempted autonomously, what fraction completed correctly without human intervention? Then ask how that number is measured. A company that can't tell you its autonomous-completion rate, broken down by task type and customer, doesn't have observability into its own product, which is itself a finding.

Probe the failure modes. What happens when the agent encounters an input it wasn't built for? Does it fail loudly, escalate to a human, or confidently produce a wrong answer? In a per-outcome business, silent wrong answers are the worst case: the customer pays, the agent fails invisibly, and trust erodes until the contract isn't renewed. Look for an evaluation harness, regression testing on prompts and model upgrades, and a real on-call function. Anthropic's own guidance on building reliable, well-scoped agentic systems is a useful benchmark, well-run agent companies tend to mirror those practices internally.

Finally, ask about the "human-in-the-loop tax." Many agents that look autonomous in the demo are quietly backstopped by human reviewers in production. That's not disqualifying, but it is a margin and scalability question. If 30% of tasks route to a human, the company is partly a BPO with software margins on only part of the volume, and your model needs to reflect that.

Model Dependency and Supply-Chain Risk

Almost every GaaS company is built on top of one or more foundation models it does not own. That dependency is the single most important supply-chain risk in the business, and it cuts several ways.

Pricing exposure is the obvious one: if the underlying provider changes per-token pricing or rate limits, the company's COGS moves overnight. Ask whether the company has committed-use discounts, multi-provider routing, or the ability to swap models without a quality cliff. A company locked to a single model with no fallback is a price-taker.

Capability and deprecation risk is subtler. When a provider ships a new model, the company's carefully tuned prompts and evals may break, requiring a re-validation scramble. Conversely, when a provider deprecates an older model, the company may be forced to migrate on the provider's timeline, not its own. And there's the existential version: the model provider could launch a competing first-party agent in the same vertical, turning a supplier into a competitor. Strategic-investor dynamics, where model labs fund the very ecosystem they might later compete with, make this worth mapping explicitly during diligence.

The defensible answer isn't "we'll never depend on a model." It's "here's our model-portability architecture, here's our eval suite that lets us qualify a new model in days, and here's our contractual relationship with each provider."

Security, Liability, and the Autonomy Surface

An autonomous agent with write access to a customer's systems is a security and liability surface unlike anything in traditional SaaS. The agent can take actions, send emails, move money, modify records, call external APIs, and every action is a potential incident.

Run the security workstream against the specific risks agents introduce: prompt injection (can a malicious input hijack the agent's instructions?), excessive permissions (does the agent have broader access than its tasks require?), and data exfiltration through tool use. The OWASP framework for LLM and agentic application risks is the standard reference, and a mature GaaS company should be able to walk you through how it mitigates each category. If the security team can't, that's a finding that should affect price.

Then there's liability. When an autonomous agent makes a costly mistake, wires money to the wrong account, gives a customer's customer bad advice, who is on the hook? Read the customer contracts. Look at the indemnification and limitation-of-liability terms. A company that has accepted broad liability for autonomous actions across a large customer base is carrying a tail risk that may not be reserved for. A company that has pushed all liability onto customers may face that as a sales objection that caps its enterprise TAM. Either way, you need to understand the exposure before you underwrite the valuation.

The Moat Audit: What's Actually Defensible

Late-stage valuations assume durable competitive advantage. For agent companies, the moat is frequently overstated in the deck and under-examined in diligence.

The model is not the moat, it's rented, and competitors rent the same one. The agent's clever prompting is not the moat, it's reverse-engineerable. The real, durable moats in GaaS tend to be three things: proprietary data and feedback loops the agent improves on (data the company captures from doing the work that no competitor can replicate), deep workflow integration that raises switching costs, and accumulated trust in regulated or high-stakes verticals where a track record is itself the product.

Ask the company to articulate, specifically, what a well-funded competitor would have to replicate to win an existing customer. If the honest answer is "a few months of engineering," you're looking at a feature, not a company. If the answer involves years of proprietary outcome data, regulatory clearances, and integration depth that would force the customer through a painful migration, the moat is real. Score it accordingly.

Org and Founder Diligence for Autonomous Products

The team that builds an autonomous product needs a particular shape. You want strong ML/eval engineering, yes, but also the operational discipline to run a product that takes real-world actions 24/7. Ask who owns reliability. Ask what happens at 3 a.m. when the agent starts failing. The answer reveals whether this is a research project wearing a company's clothes or an operationally mature business.

Watch for "agentwashing" at the org level, a team that markets autonomy it hasn't actually built, where the demo is heavily scaffolded and production is mostly human. The tell is usually a mismatch between the marketing language and the actual autonomous-completion rate. Cross-check the founder's claims against what the engineering org can demonstrate live, on a task the company didn't pre-select.

McKinsey's research on scaling AI agents in the enterprise is consistent on one point: the bottleneck is rarely the model, and almost always the surrounding organization, the evals, the guardrails, the change management, the operational muscle. Diligence the org with that lens.

The Checklist, Condensed

For the partner who wants the one-pager to take into the IC meeting:

  1. Revenue quality, Decompose by pricing model; demand consumption cohorts; reconcile booked outcomes against customer systems; flag pilot-heavy and concentrated revenue.
  2. True gross margin, Force inference, tool/API fees, vector queries, and human review into COGS; examine the margin trend, not the snapshot.
  3. Reliability, Get the honest autonomous-completion rate by task type; understand failure modes and the human-in-the-loop tax.
  4. Model dependency, Map pricing, deprecation, and competitive risk from foundation-model providers; verify model portability.
  5. Security and liability, Test against agent-specific threats; read the indemnification terms; reserve for tail risk.
  6. Moat, Identify what a funded competitor would actually have to replicate; distinguish data/integration/trust moats from clever prompting.
  7. Org maturity, Confirm operational ownership of a 24/7 autonomous product; screen for agentwashing.

Insights Most People Overlook

The best agent companies look temporarily worse on gross margin. A company aggressively investing in harder, higher-value autonomous tasks may show compressing margins today while building a deeper moat. A company with suspiciously perfect margins may simply be cherry-picking easy tasks that a competitor will commoditize. Don't reflexively reward the cleaner-looking P&L; understand why the margin looks the way it does.

Outcome-based revenue can be too good, and that's a risk, not a feature. When an agent is dramatically cheaper than the human labor it replaces, the customer eventually notices the gap between what they pay and what it costs the vendor to deliver. That gap invites in-housing, renegotiation, or a competitor undercutting on price. Wide outcome-pricing spreads aren't a durable moat; they're an arbitrage that buyers close over time. Probe how the company defends pricing as the market matures.

Falling model prices are not unambiguously good news. Investors assume cheaper inference helps GaaS companies, and on margin it does. But cheaper, more capable models also lower the barrier for new entrants and for customers to build in-house. The same cost curve that improves the income statement erodes the moat. The companies that win are the ones whose defensibility lives above the model layer.

The "we're model-agnostic" claim deserves a live test, not a slide. Many companies say they can swap models; far fewer can actually do it without a quality regression. Ask them to demonstrate the same task on two different underlying models, live, and compare outputs and latency. The gap between the claim and the demo is one of the most predictive findings you'll get.

Reliability data is a negotiating tool, not just a risk flag. If a company can't produce a clean autonomous-completion rate, that's not only a diligence concern, it's leverage. The absence of basic observability is a structural weakness that justifies a lower entry price or stronger protective terms. Use it.

References

#agent unit economics

More in Market