How to Vet an AI Agency Demo Before You Sign
A demo proves an AI agent works on the inputs the agency chose. Here are the questions and tests founders should run before signing with an AI agency.
TL;DR
An AI demo only proves the system works on inputs the agency picked. Before signing, bring your own messy data, ask to see the failures, ask what the system does when the model is wrong, and get the eval set written into the contract.

An AI agency demo proves exactly one thing: the system works on the inputs the agency chose to show you. It does not prove it works on your data, your edge cases, or your users at 2 a.m. Before you sign, you need to replace their inputs with yours, look at the failures instead of the successes, and get the test set written into the contract.
This matters more for AI projects than for ordinary software. A CRUD app either saves the record or it does not. An LLM-powered agent can look perfect on fifty curated examples and then confidently produce garbage on the fifty-first. A founder who cannot tell a rigged demo from a robust system is buying on faith.
Why AI demos are so easy to fake, even accidentally
Nobody has to be dishonest for a demo to mislead you. Three things happen naturally:
- The examples are cherry-picked. The team built the prototype against a handful of inputs and those are the ones that work. They show you what they tested. If the demo set is ten clean, well-formatted records and your real inbox is support tickets in two languages with no punctuation, you are looking at a different product than the one you will be paying for.
- Failures get filtered out before you see them. A recent post making the rounds described a benchmark that scored a model a perfect 1.00 because the harness silently skipped the three questions the model failed. The same thing happens in demos: a retry loop, a fallback to a canned answer, or a quietly dropped row hides the misses. A score of "98 percent" with no denominator and no list of the misses is the same trick in a nicer font.
- The system is tuned for the demo, not for production. Low temperature, a huge context window stuffed with your exact documents, a human pre-cleaning the input. None of that survives contact with real traffic and real costs. The stuffed context in particular is a tell: it makes every answer look grounded while quietly multiplying the per-input token bill you will ask about later.
None of this is malicious. But the burden is on you to see past it.
Bring your own inputs
The single most useful thing you can do is show up with your own data. Not a description of your data, the actual records.
Pull 20 to 50 real examples from the workflow the agent will handle. Deliberately include the ugly ones:
- A support ticket written in two languages with no punctuation
- A scanned PDF where the table is slightly rotated
- A lead form where the company field says "n/a" and the phone number is in the email field
- Two records that look like duplicates but are not
- A message that asks for something the agent should refuse or escalate
Then ask the agency to run these live, in front of you, and show you every output. Not an accuracy percentage. The raw outputs, side by side with the inputs.
If they ask for a week to "prepare the environment" first, that is useful information too. A system that needs a week of prep to handle 30 of your records is a system that will need a week of prep for every new customer segment you add.
Ask to see the failures
Flip the usual demo script. Instead of "show me it working", ask:
- "Show me the last five inputs where it got the wrong answer." A serious team has these. They keep a failure log because that is how you improve an LLM system. A team with no failures to show has either not tested or is not telling you.
- "What does the system do when the model is wrong?" The honest answer involves validation, confidence thresholds, and a route to a human. The worrying answer is "it rarely is". Every model is wrong some percentage of the time; the design question is what happens next. We covered the patterns in deterministic guardrails for AI features.
- "How do you know it has not regressed after you change a prompt?" You are listening for the word evals: a fixed set of inputs with expected outputs that runs on every change. If they do not have one, you will be the eval set, in production.
Ask what is actually running
A demo hides architecture, and architecture is what you are paying for over the next two years. Ask them to walk through the request path for one input, start to finish:
- What is the model, and is it a provider API or self-hosted?
- How many model calls per input, and how long does the slowest one take?
- What retrieval is happening, over what documents, refreshed how often?
- What is deterministic code and what is left to the model?
- What does one input cost in tokens, and what does that look like at your expected monthly volume?
Be especially alert to complexity theater. Agencies sometimes present a five-agent pipeline with planners, critics and orchestrators because it looks sophisticated on a slide. In our own outreach tooling, where we scrape each prospect's site and have a local model draft one tailored email, the version that extracted facts and wrote the draft in a single call beat the multi-step chain on both cost and output quality. Every extra hop was another place for the pipeline to lose context or hallucinate a bridge between steps. We wrote up the trade-offs in single call vs agent chains.
Simple is not automatically better, but a team that can explain why each component exists is a team that understands the system. A team that cannot is maintaining something they inherited from a tutorial.
Check the production story, not just the prototype
The gap between a demo and a deployed system is where most AI projects die. The good news is that the 30 records you brought and the request walkthrough you just sat through give you everything you need to make these questions concrete instead of hypothetical:
- Who owns the prompts and the eval set after launch? You should. The eval set is your 30 records plus every failure they showed you, with agreed correct outputs. Make sure it is named in the deliverables, not described as "documentation".
- What happens when the model provider deprecates the version you tested on? Providers retire models on a schedule. The answer you want is "we swap the model, re-run the acceptance set, and show you the diff". If re-running the set is not a one-command operation, ask why not.
- How is cost capped? Take the per-input token cost from the walkthrough, multiply by your monthly volume, and ask what stops the bill from being ten times that. A runaway retry loop or a spam burst against an unprotected endpoint will find out for you. Rate limits and per-user budgets belong in the build, not in a later phase.
- What is monitored? At minimum: latency, error rate, fallback rate, and some sampled quality review. Fallback rate is the one to watch. It is the production version of "what happens when the model is wrong", and if it climbs from 5 percent to 20 percent after a prompt change, that is your regression alarm. If their answer is "we'll add logging", they have not shipped one of these before.
- What does month three look like? New edge cases arrive, the eval set grows, a prompt gets tuned, a model gets updated. Who re-runs the set each time, and what does that cost per month?
A team that has taken agents from pilot to production will answer these without hesitation, because they have been burned by each one. Our pilot to production guide lists the specific failure modes we see most often.
Put the test set in the contract
Here is the move that separates a hopeful engagement from a safe one. Take the inputs you brought to the demo, agree on what a correct output looks like for each, and write that into the statement of work as the acceptance test.
This does three things at once. It forces both sides to define "working" in concrete terms before money changes hands. It gives the agency a target they can actually hit, which good agencies appreciate. And it gives you a regression suite you will keep running for the life of the product.
Agree on a threshold, not perfection. Something like "at least 90 percent of the acceptance set correct, zero outputs in the unsafe category, and every incorrect output routed to review" is realistic for most document and sales workflows. A hundred percent is a sign nobody has thought about the failure path.
A short checklist for the meeting
Walk in with this and you will learn more in an hour than in three polished pitches:
- Run my 30 real inputs live, show every output
- Show me five recent failures and what you changed because of them
- Walk me through one request end to end: models, calls, retrieval, cost
- What happens when the model is wrong?
- Who owns prompts, evals and monitoring after launch, and what does month three cost?
- Will you put the acceptance set and threshold in the contract?
An agency that welcomes these questions is one you can work with. One that treats them as hostile is telling you how the project will go.
If you are evaluating AI partners right now and want a second opinion on a proposal, or want to see how we would run your 30 ugliest records, let's talk.
Frequently asked questions
Why can't I trust a polished AI agency demo?
Because the agency chose the inputs. LLM systems look flawless on curated examples and fall apart on edge cases, so a demo tells you the happy path works, not that the product will survive your real users.
What should I bring to an AI agency demo?
Twenty to fifty real examples from your business, including ugly ones: half-filled forms, mixed-language messages, duplicate records, scanned PDFs. Ask the agency to run them live and show you every output, not a summary score.
How do I compare two AI agencies after demos?
Run both on the same set of your real inputs and compare the failures, not the successes. Then ask each one how they would detect and handle those failures in production and what the ongoing cost looks like.
Is a simpler architecture a red flag in an AI agency?
Usually the opposite. A single well-prompted model call with strict validation often beats a multi-agent chain on cost and accuracy. Be more suspicious of complexity you cannot test than of simplicity you can.
How Pykero can help
Services related to this article.
Building something like this?
Pykero Agency designs and ships production web, mobile, SaaS, and AI products.
Talk to us →

