How to Run a Procurement Process for Buying AI Agents (Without Getting Burned)
Buying an agent is not buying software, and it is not hiring a contractor either. It sits in an awkward middle, which is exactly why most procurement teams stumble on their first one. This guide walks the full process: scoping the outcome instead of the feature list, evaluating Agentic AI-as-a-Service (GaaS) vendors on reliability and security rather than demo polish, negotiating per-task and per-outcome pricing, and writing contract terms that actually protect you when an autonomous system makes a costly mistake. The short version: treat the agent like a probationary employee you're also indemnifying, and write the contract accordingly.
Table of Contents
- Why Buying an Agent Breaks Your Normal Procurement Playbook
- Stage One: Define the Outcome, Not the Tool
- Stage Two: Build the Vendor Shortlist
- Stage Three: The Evaluation That Actually Predicts Production
- Reliability and the Eval Harness
- Security and Data Boundaries
- Integration Reality Check
- Stage Four: Pricing Models and How to Read Them
- Stage Five: The Contract Terms That Matter
- Who Should Be in the Room
- A Realistic Timeline
- Insights Most People Overlook
- References
Why Buying an Agent Breaks Your Normal Procurement Playbook
Most procurement functions have two well-worn templates. One is for software: you compare feature matrices, check the SOC 2 report, negotiate per-seat pricing, and sign a multi-year deal. The other is for services: you scope a statement of work, vet the firm's track record, and pay for deliverables or hours.
An AI agent sold as a service is neither, and pretending it's one of the two is how you end up with a contract that doesn't fit the thing you bought. A vertical agent that resolves customer tickets, reconciles invoices, or qualifies leads is software in the sense that it's code running on someone's infrastructure. But it behaves like a worker: it takes actions in your systems, it makes judgment calls, it occasionally gets things wrong, and its quality drifts as the underlying models and your data change. You are buying probabilistic labor with a license agreement stapled to it.
That mismatch shows up everywhere. Per-seat pricing makes no sense for something that has no seats. Feature checklists tell you what the demo can do, not whether the agent will hold up across 10,000 real tickets. And the standard software warranty disclaimer ("provided as-is, no liability for indirect damages") is genuinely dangerous when the product can autonomously issue a refund, send an email to your customer, or push a change to a production system. The procurement process has to be rebuilt around those differences. The rest of this piece is that rebuild, stage by stage.
Stage One: Define the Outcome, Not the Tool
The single most common procurement mistake in GaaS is writing a requirements document that describes an agent instead of a result. "We need an AI agent that can read emails, classify them, and draft replies" is a feature list. It invites every vendor to nod along and demo exactly those three things, and it tells you nothing about whether the thing works.
Flip it. Write down the outcome you're paying for, the volume, and the bar for "good enough." Something like: "Resolve 60% of tier-1 support tickets end to end, with a customer-satisfaction score no lower than our human baseline of 4.2, at a cost per resolved ticket under $1.10." Now you have something measurable. Every vendor conversation, every pilot, every contract clause can be tied back to that line.
This is also where you decide what the agent is allowed to do on its own versus where a human stays in the loop. An agent that drafts replies for human review is a fundamentally different risk profile, and a different price, than one that sends replies autonomously. Get explicit about the autonomy boundary now, because it drives both your security review and your liability exposure later. If you're early in your agent journey, mapping this against an enterprise agent-adoption maturity model helps you avoid biting off more autonomy than your operations can supervise.
One more thing worth doing here: write down your current cost and quality baseline before you talk to a single vendor. Without it, you cannot judge ROI, and you will be at the mercy of whatever benchmark the vendor brings. McKinsey's research on enterprise AI value capture keeps landing on the same point, the organizations that get returns are the ones that tie AI initiatives to specific business outcomes and measure them rigorously, rather than buying capability and hoping value follows.
Stage Two: Build the Vendor Shortlist
The GaaS market is loud and young. You'll find three rough categories of seller, and they're not interchangeable.
First, the vertical specialists, companies that do one job deeply (claims processing, SDR outreach, code review, revenue-cycle management for healthcare). They tend to have better domain accuracy and pre-built integrations for their niche, but they lock you into that niche.
Second, the horizontal platforms that let you build or configure agents across many functions. More flexible, but you're often buying a toolkit and doing more of the assembly yourself, which shifts work onto your internal team.
Third, the incumbents bolting agents onto existing suites, your CRM, ERP, or service-desk vendor now has an "agent" SKU. Convenient integration, but the agent is often less capable than a focused startup's, and the pricing is bundled in ways that hide the real cost.
Sourcing tip that saves weeks: ask every shortlisted vendor for a reference customer running the same use case at similar volume, and actually call them. Ask the reference what broke, what the real accuracy was after three months (not at launch), and how responsive the vendor was when the agent did something dumb. The gap between launch-day accuracy and month-three accuracy is where most disappointment lives. Andreessen Horowitz's writing on the emerging economics of AI agents is a useful primer on how these business models actually work before you sit across the table from a founder pitching you.
Stage Three: The Evaluation That Actually Predicts Production
A polished demo predicts almost nothing. Demos are run on cherry-picked inputs by people who know exactly how the agent behaves. Your job is to find out how it behaves on your messy, real-world inputs at scale. This is the part of the process that most teams under-invest in, and it's the part that determines whether you join the majority of pilots that never reach production.
Reliability and the Eval Harness
Insist on a paid pilot run against your own data and your own edge cases, scored against the outcome metric you defined in stage one. Build an evaluation set of real cases, including the weird ones, the angry customers, the malformed invoices, the ambiguous requests. Run the agent on all of them and measure not just average accuracy but the shape of the failure distribution. A 92% success rate sounds great until you learn the 8% failures are confidently wrong and silently expensive.
Two numbers matter more than headline accuracy. One is the escalation rate, how often the agent correctly recognizes it's out of its depth and hands off, versus barreling ahead. An agent that fails loudly is far safer than one that fails quietly. The second is consistency, run the same input ten times and see if you get ten similar answers. Agents are non-deterministic by nature, and a vendor who can't show you how they constrain that variance hasn't solved it.
Security and Data Boundaries
An agent that takes actions in your systems is an attack surface and a data-exfiltration risk in one. Your security review needs to go beyond the standard SaaS questionnaire. Ask where your data goes, whether it's used to train shared models, which third-party model providers sit behind the agent, and what happens to your data inside those providers. Ask specifically about prompt injection, whether a malicious customer email or a poisoned document can hijack the agent into taking unauthorized actions. The OWASP project's Top 10 for LLM Applications is the right checklist to bring into this conversation; if the vendor hasn't heard of it, that's a finding in itself.
Pin down the permissions model too. The agent should operate under scoped, revocable credentials with least-privilege access, not a borrowed admin account. You want to be able to see every action it took, attribute it, and pull the plug instantly. If the answer to "show me the audit log of everything this agent did last Tuesday" is vague, walk.
Integration Reality Check
The agent has to touch your systems to do anything useful, and that integration is almost always harder than the demo suggests. Legacy systems, weird auth, rate limits, and data that doesn't match the agent's assumptions are where timelines slip. During the pilot, connect the agent to a real (or realistic staging) version of your actual stack, not a clean sandbox. The integration burden deserves its own line in your evaluation, because a brilliant agent that can't reliably reach your data is worthless.
Stage Four: Pricing Models and How to Read Them
GaaS pricing is genuinely new, and the model the vendor offers tells you how confident they are in their own product.
- Per-seat or flat subscription. Familiar and predictable, but it decouples price from value. You pay the same whether the agent does great work or sits idle. Vendors like it because revenue is stable; you should be wary because it puts none of the performance risk on them.
- Per-task / per-action. You pay each time the agent does a unit of work, per ticket handled, per document processed, per call made. This scales with usage and is easy to model, but watch the definition of "task." A vendor can pad volume by counting retries, sub-steps, or low-value actions as billable tasks.
- Per-outcome. You pay only when the agent achieves the defined result, a resolved ticket, a qualified lead, a booked meeting. This aligns incentives beautifully and is the model serious operators increasingly prefer, because the vendor only wins when you do. The catch is that "outcome" must be defined with lawyerly precision, and you need a shared, trusted way to measure it. A vendor confident enough to price per outcome is also a vendor confident in their reliability, which is its own useful signal.
Whatever model you choose, build a realistic total-cost picture, not just the headline rate. Per-task pricing at $0.40 a task is cheap until you multiply by your real volume and add the human oversight, integration maintenance, and escalation handling that no pricing page mentions. The fully loaded cost is what your CFO will hold you to.
Stage Five: The Contract Terms That Matter
This is where the agent's worker-like nature collides with software contract templates, and where you earn your keep. A few terms deserve fights.
Service levels tied to your outcome metric. The SLA shouldn't just promise uptime; it should commit to accuracy or resolution thresholds, with credits or termination rights if the agent drifts below them. Model drift is real, an agent that performed at 90% in the pilot can degrade as upstream models change. Your contract needs teeth for that scenario.
Liability and indemnification for autonomous actions. If the agent autonomously issues a wrong refund, sends a defamatory email, or makes a decision that lands you in regulatory trouble, who pays? The default software disclaimer says "not us." For an autonomous actor, that's unacceptable. Negotiate explicit liability for the agent's actions within its defined authority, and carve out indemnification for the categories that scare your legal team most.
Data rights and exit. Get clear language that your data isn't used to train models that benefit competitors, and that you can extract your data, configurations, and the institutional knowledge the agent has accumulated when you leave. Switching costs in GaaS can be brutal precisely because the agent learns your business; don't let that learning become a hostage.
The kill switch and human override. Contractually guarantee your ability to pause or constrain the agent instantly, and to set autonomy boundaries that the vendor can't override via a silent update. A pushed update that quietly expands what the agent does on its own is a governance nightmare waiting to happen. For a fuller catalog of contract and demo warning signs, the dedicated piece on procurement red flags when buying agents is worth reading alongside this one.
Who Should Be in the Room
Buying an agent is a cross-functional act, and the most common org failure is letting one department run the whole thing. IT alone buys for integration and ignores business outcomes. The business unit alone buys on demo dazzle and ignores security. You want, at minimum: the business owner who owns the outcome, IT/engineering for integration and the permissions model, security and data governance, legal for the liability terms, and finance for the TCO model. The friction between IT and the business over who owns the agent is real and worth getting ahead of before it stalls the deal.
A useful pattern emerging in more mature organizations is a small central group, an agent center of excellence or an early AgentOps function, that owns the procurement playbook so every business unit isn't reinventing it. Even if you're buying your first agent, writing down what you learn turns this one painful process into a reusable template.
A Realistic Timeline
For a meaningful enterprise agent, plan on roughly eight to fourteen weeks from kickoff to signed contract: a week or two to define the outcome and baseline, two to three weeks of sourcing and reference calls, four to six weeks for a paid pilot against real data (the part you cannot rush), and two to four weeks of security review and contract negotiation running partly in parallel. Teams that compress the pilot to a two-day demo are the same teams writing post-mortems six months later. The pilot is not the expensive part of this process, a bad production deployment is.
Insights Most People Overlook
The vendor's willingness to price on outcomes is the single best signal you have. Forget the demo. A vendor who will tie their revenue to your resolved tickets is telling you, with their own money, that they trust their agent's reliability. A vendor who insists on flat subscription pricing for an outcome-driven product is hedging against their own failure rate. Make them explain why.
You're underwriting model risk you didn't choose. Most agents sit on top of third-party foundation models the vendor doesn't control. When that model gets updated, deprecated, or quietly swapped for a cheaper one, your agent's behavior can change overnight, and you'll find out from your customers, not your vendor. Ask explicitly which models power the agent, what the vendor's policy is on model changes, and whether they re-run their evals when the underlying model shifts. Almost nobody asks this, and it's the thing most likely to break a working agent.
The cheapest agent often has the highest total cost. A low per-task rate frequently correlates with a higher error rate, which means more human escalation, more cleanup, and more reputational risk. The fully loaded cost per correct outcome, not per task, is the number that matters, and it sometimes inverts the price ranking entirely. The "cheap" vendor can be the expensive one once you count the humans cleaning up after it.
Procurement doesn't end at signing, agents need ongoing supervision the contract should fund. Unlike software you buy and forget, an agent needs monitoring, periodic re-evaluation, and feedback loops to stay good. Build the expectation (and the budget) for continuous oversight into the deal from day one. Teams that treat agent procurement as a one-time transaction are the ones surprised when quality erodes in month four.
A successful pilot can lie to you about scale. An agent that nails a controlled 500-case pilot can behave very differently at 50,000 cases a month, where rare edge cases stop being rare in absolute terms and integration rate limits start to bite. Insist on a volume ramp clause and watch the metrics as you scale, rather than assuming pilot performance is a flat line you can extend.
References
More in Adoption
- Building an Internal Agent Center of Excellence: The Org Muscle That Decides Whether Agents Stick
- Why IT and the Business Fight Over Who Owns the Agents
- Who Owns the Agents Inside a Company? The Accountability Question Nobody Asked Until It Broke
- Onboard Your AI Agent the Way You'd Onboard a New Hire (Not the Way You Install Software)
- AgentOps Is Becoming a Real Job, Here's What That Function Actually Does