The problem with how this decision is usually made
Every guide to selecting an AI partner says the same four things: look for deep expertise, relevant industry experience, a proven track record, and clear communication. Those criteria describe every vendor who has a website. They disqualify nobody, which means they are not criteria at all — they are a ritual that lets a committee feel it has done diligence.
Vetting works the other way round. You are not looking for reasons to say yes; you are looking for the one reason that makes everything else irrelevant. This article is a list of those reasons.
The disqualifying signal in an AI consultant is not a weak portfolio. It is a refusal to name where the approach fails — because someone who has shipped has a list of failure modes, and someone who has only demoed does not.
Why the failure question is the decisive one
Ask a vendor to describe a project that did not work out, and watch what happens to the conversation.
Someone who has delivered will tell you about a specific engagement: what the client believed at the start, what turned out to be true, what they got wrong, what it cost, and what they changed in their own method as a result. The story will be slightly unflattering. They will not enjoy telling it. That reluctance is itself a good sign, because it means the memory is real.
Someone who has demoed will do one of three things. They will say every project succeeded and attribute any difficulty to client-side factors. They will reframe the question as an opportunity to describe their process. Or they will name a failure so small and so abstract ("we had a scope conversation that went long") that it functions as another success story.
All three answers tell you the same thing: there is no accumulated failure knowledge behind the proposal. And that matters because the failure modes in this work are not exotic. They are specific, repetitive and known: the process turns out to have an undocumented exception path; the data exists but not in a queryable shape; the people who own the process leave; the model is accurate on the cases the client thought of and wrong on the ones they did not. A vendor who has shipped knows this list and will volunteer it before you ask.
Eight disqualifying signals
The table is the short version. Below it, the four that buyers most often misread.
| Signal | What it looks like | Why it disqualifies |
|---|---|---|
| Every past project succeeded | No specific failure named, or the failure is a client behaviour | There is no failure knowledge to apply to your project |
| No written scope before work starts | "We'll figure out the details as we go" | The scope is the deliverable that makes every later step possible |
| Pricing is hourly | A rate card, no outcome framing | Hourly pricing makes the correct engineering decision financially irrational for the vendor |
| No stated boundary on what they do not do | "We do whatever the client needs" | A practice with no boundary has no position, and will accept work it cannot do |
| You cannot see their own software | Products claimed but never linked or shown | Claimed products are the easiest thing to verify and the most commonly faked |
| They answer every question confidently | No "it depends", no conditions, no "I would need to look" | Certainty across every domain is a sales posture, not a capability |
| The proposal is a technology choice | A stack named before the problem is stated | The technology is a consequence of the diagnosis, never the diagnosis |
| All the risk is yours | Success fee, no exit, or a scope that only grows | A vendor who carries no delivery risk has no reason to be right |
Signal 1: "every project succeeded" is the strongest negative
This is the one buyers most often read backwards. A long success record is treated as the headline credential, so a vendor with a perfect record looks like the safest choice.
Consider what generates such a record. In this work, some engagements should be refused. The process cannot be described, the data is not there, or the honest answer is that a deterministic rule solves it and AI would be a downgrade. A practice that never refuses anything has therefore never exercised judgment — it has accepted every engagement and reported on the ones that went well.
There is a second reason to distrust it. The failure rate people quote for AI pilots is high enough that a perfect personal record is statistically unlikely. I am not going to cite a figure, because the numbers that circulate are not reproducible across process types and I have not run the study. The point stands without one: if a vendor's record has no failures in it, either the sample is small, the definition of failure is generous, or the record is being curated for you.
What replaces it: ask for the failure list first. A vendor who leads with three things that went wrong, unprompted, has told you more about competence than a reference list will.
Signal 2: no written scope, and the reason is usually stated as flexibility
"We prefer to stay agile" sounds like an engineering virtue. In an AI engagement it is a specific risk, because the capability boundary of a model-based system is invisible in a way that a normal software boundary is not. Nobody can tell by looking at a document-parsing workflow where accuracy falls off. So scope creep in this work does not look like a client asking for extra features — it looks like a system that is 94% right on the cases that matter and nobody having agreed in advance what happens to the other 6%.
The useful artefact is not a feature list. It is a written statement of what the system must not do, and what happens when it cannot decide. A vendor who will not commit that to paper before starting has not thought about it, or has thought about it and prefers not to be held to it.
Signal 3: no boundary, and the "we do everything" answer
Ask directly: what kind of work do you turn down?
An answer that says everything is available tells you the practice has no thesis. Three concrete refusals are worth more than twenty capabilities:
- Work where the process cannot be described by the person who owns it, because there is nothing to automate yet.
- Work inside a fixed architecture the vendor cannot influence, because implementing someone else's design is a different and cheaper service.
- Work where the honest answer is that no model is needed, because there is a rule, a constraint, or a scheduled job that solves it for nothing.
A vendor with a boundary will sometimes tell you that you are not a fit. That is the interaction you should be paying for, and it is the one vendors optimise against, because it feels like losing a deal. If you want to see what the boundary looks like stated in advance rather than discovered on a call, the engagement model and its four refusals are published rather than negotiated.
Signal 4: claimed products you cannot verify
Almost every independent practitioner will claim they have shipped something. This is the single easiest claim to check and the one most often left vague.
Look for a name, a link, and something that resolves. Then look at whether the thing behaves like a product: does it have a pricing page, a changelog, a status page, users who talk about it. A practitioner who genuinely operates software has a different relationship to it than one who built a prototype in 2023 — they will talk about its running costs, its failure modes, and the things they would rebuild.
The reason this matters for your project is not the product itself. It is that a person who operates something they built has been on the receiving end of their own architectural decisions for years. That changes what they propose to you, and it changes what they refuse to propose.
What to ask, in order
A single hour of questions, in this sequence, separates the two populations faster than a portfolio review:
- Tell me about a project that failed. Listen for the specifics: the belief at the start, what was actually true, the cost, the change to their method.
- What would make you turn this down? A vendor with no answer will accept the engagement regardless of fit.
- Write down what this system must not do. If the answer is that this comes later, the boundary will be discovered in production.
- Show me your own software. Named, linked, and currently running.
- Where does this approach break? The failure mode should arrive without hesitation and should be specific to your kind of process.
- What is the first paid step, and what do I keep if I stop there? The answer should be small, fixed, and leave you holding a document.
When this advice is wrong
Two honest caveats, because a list of disqualifying signals is itself a heuristic and heuristics fail.
First, a large established consultancy will often fail signals two and three — it has a standardised process and a broad service catalogue by design, and for a large programme with regulatory exposure that standardisation is the value you are buying. This list is calibrated for the engagement where one person or a small team carries the architecture, which is a different purchase.
Second, some of the answers you are screening for require trust to be useful. If a vendor tells you a failure story, you have no way to verify it from outside. The signal is not the verifiable fact; it is the shape of the answer. Someone who has shipped produces a story with inconvenient details. Someone who has not produces a lesson.
Finally, the uncomfortable part: this list applies to me. Ask me the six questions in order and you will get specific answers, including the refusals and the failures. That is the only version of this article worth publishing — a vetting guide whose author is exempt from it is a sales document wearing a checklist.
The next step
Before your next vendor call, write down the one answer that would make you walk away. Do it in advance, so you are not constructing the standard while listening to someone persuasive.
Then ask the six questions and write down what you heard, in their words, immediately afterwards. Two vendors later, comparing the notes will tell you more than comparing the proposals.
Keep reading
- What a Real AI Engagement Looks Like, Week by Week2026-03-317 minAI Strategy
- An AI Readiness Assessment Should Produce a Process Register, Not a Maturity Score2026-03-187 minAI Strategy
- The 90-Day AI Roadmap: Ship One Workflow and Label Everything Else a Guess2026-03-168 minAI Strategy
- AI Adoption Without Layoffs Is a Reallocation Decision, Not a Kindness2026-02-228 minAI Strategy
- AI Competitive Advantage for a Small Company: Win on Weeks, Not Budget2026-01-257 minAI Strategy