BusinessMCP

AI agents

How to Evaluate AI Agents: A Buyer’s Framework for 2026

Gartner estimates only about 130 of the thousands of “agentic AI” vendors are real. Here is the buyer-side framework we use to review agents for our directory: the agent-washing test, a weighted evaluation rubric, the questions that make weak vendors visibly uncomfortable, and a pilot design with a kill threshold.

By the BusinessMCP team12 min readAugust 15, 2026
How to Evaluate AI Agents: A Buyer’s Framework for 2026 — illustrated overview

Key takeaways

  • Agent washing is real: Gartner estimates only about 130 of thousands of “agentic AI” vendors are genuinely agentic — but a well-built workflow can still be worth buying, so evaluate the outcome, not the label.
  • Score vendors on a weighted rubric — data access, guardrails, observability, model flexibility, security, pricing transparency, proof and exit cost — instead of comparing demos.
  • The highest-signal vendor questions are about failure: what happens when the agent is wrong, what gets logged, and what a refund-worthy miss looks like.
  • A pilot without a predefined kill threshold is a soft launch — decide the success metric, the baseline and the walk-away number before day one.

Why evaluating AI agents is genuinely hard right now

Learning how to evaluate AI agents matters more in 2026 than in any prior software cycle, for one reason: the gap between claim and capability has never been wider. Gartner estimates that of the thousands of vendors marketing “agentic AI,” only about 130 are real — the rest rebrand chatbots, RPA and assistants. The same release predicts over 40% of agentic AI projects will be canceled by end of 2027 on costs, unclear value or inadequate risk controls.

We review agents for our directory of ~150 tools, so we run some version of this evaluation weekly. This guide is the framework: an agent-washing test, a weighted rubric, vendor questions and a pilot design. It assumes you already know what an agent is — if not, start with our explainer and the use-case maturity map, then come back.

The agent-washing test: is it actually an agent?

Anthropic’s engineering team draws the cleanest public line in Building Effective Agents: workflows orchestrate LLMs through predefined code paths, while agents dynamically direct their own process and tool use. Most products marketed as agents are workflows with an LLM step — a fixed pipeline that calls a model once or twice.

  • Does it decide, or just execute? Give it an ambiguous goal in the demo. An agent chooses tools and sequencing; a workflow asks you to configure the sequence.
  • Can it recover? Feed it a failure (a bounced email, a missing field). Agents replan; workflows error out or silently skip.
  • Does the tool list matter? Ask what tools it can call and whether it selects among them at run time. A one-tool “agent” is an API wrapper.
  • Is there a loop? Agents observe results and iterate; single-pass generation — however good — is not agency.

The evaluation rubric we score against

Demos optimize for wow; rubrics optimize for regret-avoidance. This is the weighted scorecard we apply — adjust the weights to your risk profile, but score every vendor on all eight lines so the comparison is honest:

AI agent evaluation rubric (score each criterion 1–5, multiply by weight)
CriterionWeightWhat a 5 looks like
Data & tool access20%Connects to your real systems (CRM, analytics, inbox) via supported integrations or an open protocol — not CSV uploads
Guardrails & approvals20%Consequential actions (send, spend, write) are approval-gated by default, with autonomy you can widen deliberately
Observability15%Every action logged and inspectable: what it did, why, what it cost — auditable without vendor help
Model flexibility10%Model-agnostic or bring-your-own-key; no hard lock to one provider’s model roadmap
Security & compliance10%Clear data-handling terms, scoped permissions, tenant isolation; compliance add-ons priced upfront
Pricing transparency10%Published prices or a decomposable quote; metered units defined in writing
Evidence it works10%Reference customers you can talk to, or public numbers — not testimonial screenshots
Exit cost5%Your data, prompts and configuration export cleanly; leaving is priced into the decision

Two lines deserve their 20% weights. Data access because a context-starved agent hallucinates confidently regardless of model quality. Guardrails because the difference between an asset and an incident is whether the agent can send, spend or delete without a human checkpoint. Pricing transparency gets its own deep-dive in the pricing guide.

Questions to ask every vendor

The highest-signal questions are about failure, not features. Strong vendors answer these instantly because they have thought about them; weak ones improvise:

  1. 1What does the agent do when it is not confident? Show me an escalation, not a success.
  2. 2Show me the log of a real run — every tool call, every decision, the cost. Can my team access this ourselves?
  3. 3What is the worst thing your agent has done for a customer, and what changed afterward?
  4. 4Which actions require human approval by default, and can I widen or narrow that set?
  5. 5What exactly is the billable unit, and who audits it? (Get this in writing.)
  6. 6Which model does it run on, what happens when that model is deprecated, and can I bring my own key?
  7. 7If we leave after a year, what do we take with us?
  8. 8Can I speak to two customers at my scale who have run it for six months or more?

One meta-signal: vendors who volunteer their limitations tend to be the real ones. In our own reviews we publish cons and pricing catches for every tool — because a review with no cons is an ad.

Design a pilot that can fail

Per Gartner, unclear business value is a leading cancellation cause — and unclear value is usually baked in at pilot design, when nobody defines what failure looks like. A pilot that cannot fail is a soft launch. The structure that works:

Pick ONE metric the agent should move (resolution rate, meetings booked, hours saved)
Measure the pre-agent baseline for 2–4 weeks
Run a scoped 30–60 day pilot in assisted mode, logging everything
Compare against baseline AND against the fully-loaded cost, including your review time
Decide against thresholds you wrote down on day one: expand, renegotiate, or kill

Write the kill threshold before the pilot starts. Deciding after the fact, with sunk cost and a friendly account executive in the room, is how 40% cancellation rates happen — slowly.

The measurement half of this — cost lines, value lines and the ROI-math traps — is covered in depth in the AI agent ROI guide; for sales agents the AI SDR ROI calculator does the arithmetic for you.

Red flags and green flags

A compressed field guide from cataloging ~150 agents:

What we flag when reviewing agents
Red flagGreen flag
Demo only works on curated sandbox dataVendor offers to pilot on your real, messy data
“Fully autonomous from day one” as the pitchApproval gates on by default; autonomy is earned
No visible logs — “trust the dashboard”Per-run traces you can export and audit
Unexplainable credit systems and unpublished pricingDecomposable pricing with the metered unit in writing
Testimonial screenshots as evidenceReference customers you can actually call
Roadmap answers to capability questionsA straight “it can’t do that today”

When you are ready to shortlist, the directory organizes the field by category with review pages — pros, cons and verified pricing — on the tools we have evaluated hands-on.

Frequently asked questions

What is agent washing?

Rebranding existing software — chatbots, RPA, assistants — as “agentic AI” without substantial agentic capability. Gartner called the practice out in a June 2025 press release, estimating only about 130 of thousands of agentic AI vendors are real. The practical test: does the product dynamically decide which tools to use and recover from failures, or does it execute a predefined pipeline with an LLM step?

What criteria should I use to evaluate an AI agent?

Eight, weighted: data and tool access (20%), guardrails and approvals (20%), observability (15%), model flexibility (10%), security and compliance (10%), pricing transparency (10%), evidence it works (10%), and exit cost (5%). Score every vendor on all eight rather than comparing demos — demos are optimized to hide exactly the lines that matter.

How long should an AI agent pilot run?

Thirty to sixty days of live use, preceded by two to four weeks of baseline measurement on the metric you expect the agent to move. Shorter pilots measure novelty; longer ones without thresholds become soft launches. The non-negotiable part is writing the success and kill thresholds down before day one.

Should I buy an AI agent or a workflow tool?

Buy whichever reliably produces the outcome — a deterministic workflow is often the better choice for well-defined, repetitive processes, and its predictability is a feature. The failure mode to avoid is paying agent prices and accepting agent risk for workflow capability. Classify the product honestly first (the agent-washing test above), then judge the price against the category it really belongs to.

BM

BusinessMCP Team

Every guide is written from running BusinessMCP on its own platform — the match rates, reply rates, and deliverability lessons are from our own data, not recycled blog folklore. About BusinessMCP

Turn your business into one AI-ready MCP server

Connect your tools, install one tracking script, and expose your unified data to any AI agent through a single secure endpoint.

Get started free