You choose an AI agent development partner by testing for production evidence rather than demo quality: agents genuinely running in daily use, success defined as a number, evaluation and observability built in as defaults, cost per task engineered rather than discovered, and clear ownership of the system when the engagement ends. Any firm can show a convincing demo — the demo is the easy part. What separates partners worth hiring is what they do after it.
Why this decision is hard right now
The market for AI agent development has grown faster than the evidence base underneath it. Directories and listicles rank dozens of firms; nearly all of them show the same things — logos, testimonials, a portfolio of builds — and none of those artefacts distinguishes a partner whose agents survive production from one whose agents impressed a stakeholder meeting and were quietly switched off within six months.
The failure mode this piece exists to help you avoid is specific: most agent projects do not fail at the build. They fail after it — unmeasured, unwatched, costing more per task than anyone budgeted, drifting below a quality bar nobody defined. So the criteria that follow all test one underlying question: does this firm treat production as the job, or as the thing that happens after the job?
The seven criteria
1. Production evidence, not portfolio volume. Ask about a specific agent running in production today: how long it has been live, what it does daily, and what happened in the months after launch. One agent still running and measured a year on is stronger evidence than ten launched and unaccounted for. Be direct: “how many of the agents you have built are still in daily use?” — the comfort of the answer tells you most of what you need.
2. Success defined as a number, agreed before the build. A serious partner will insist on agreeing the measure first — hours saved per cycle, cost per task, review time, throughput — and will want to baseline your workflow before promising anything. If success is described in adjectives, or a guarantee arrives before anyone has looked at your data, the diligence has already failed. The discipline to look for is the one described in how to evaluate an AI agent: the measure is the contract.
3. An evaluation harness, not a testing phase. Ask how they know an agent is good enough to ship and how they know it stays good enough. The right answer involves a repeatable harness — the agent scored against real, representative data on every change — not a QA pass at the end. This is the single sharpest technical differentiator between firms, because it cannot be improvised at handover.
4. Observability as a default. An agent in production will fail in ways nobody predicted; the question is whether those failures are diagnosable. Ask what a trace of a single run looks like — every retrieval, tool call, and validation step, with cost and latency attached — per what to monitor in a production agent. A partner who cannot show you a trace is offering you an agent you will not be able to debug.
5. Cost per task, engineered. A demo costs pennies; production multiplies every inefficiency by thousands of runs a month, and most of an agent’s running cost is fixed by architecture long before the first invoice. Ask what their last agent cost per completed task and what budget it was designed to. A firm that has never been asked is a firm that has never had to answer.
6. Ownership and exit. Agree before the build what you hold at the end: the code, the prompts, the evaluation sets, the infrastructure, the documentation. The wrong answer is a dependency dressed as a deliverable — a system only the vendor can operate, priced accordingly forever. The test is simple: “if we parted ways in a year, what would we need from you to keep it running?”
7. The willingness to say no. Some workflows do not need an agent — they need a pipeline, a script, or RPA, at a fraction of the cost and none of the non-determinism. A partner who has never told a prospect “you don’t need an agent for this” is selling agents, not solving workflows. It is worth asking for an example directly.
The questions, in one place
Put these to every firm on your shortlist and weigh the answers against each other:
- Tell us about one agent of yours in production today. How long has it been live, and what number is it measured against?
- How do you evaluate an agent before it ships — and on every change after?
- Show us a trace of a single production run.
- What did your last agent cost per completed task, and what budget was it designed to?
- What do we own when the engagement ends?
- When did you last advise a client not to build an agent?
None of these requires technical depth to ask, and collectively they are hard to bluff. A firm that answers all six specifically is worth your shortlist regardless of its size or its position in anyone’s ranking.
Where we stand
Agent Foundry Labs is a production AI agent studio in London, and this piece is, unavoidably, also a description of our own standard — we publish the two agents we run in production, including the tender-intelligence agent we operate for Magellan Circle, with the measures attached, and we structure every engagement as Scope, Build, Run so the number is agreed before the build and reported against while the agent is live. We would rather you put the six questions above to us and to everyone else you are considering; they select for the discipline that keeps agents running, whoever you hire.
If you are evaluating partners for a specific workflow, book a 30-minute call — the first conversation is exactly the scoping test this piece describes.