A broken tool usually tells you it’s broken.
That’s the deal we’ve had with machines for about 200 years. Something fails, something makes noise. A warning light. An error message. A line that stops moving. The failure announces itself, and somebody goes and looks.
AI doesn’t work that way.
When an AI tool starts getting things wrong, it does not slow down, it does not flag anything, and it does not sound any less sure of itself. It answers in the same tone it used when it was right. Same speed. Same clean formatting. Same air of having confidence about what it tells you.
So you won’t hear about it from the tool. You’ll hear about it from a customer who noticed. Or from a board member who asked a follow-up question nobody could answer. Or — and this is the far more common outcome — you won’t hear about it at all.
Which leaves one question worth sitting with before you automate anything else: in your company, who would have caught it?
Hi, I’m Jeff Payne. You’re listening to The Jeff Payne Show, Episode #82: Confidence Isn’t a Check.
Almost every conversation happening in businesses right now about AI is a conversation about scope. What should we automate? How much of it? How fast. Who owns it?
Those are reasonable questions. They’re also the wrong first question, and they’re wrong because they all assume the constraint is capability—what the tool can do.
The constraint isn’t what the tool can do. The constraint is what your team can check.
Every AI output that enters your business arrives as a claim. A number, a recommendation, a draft, a summary. And a claim is only worth what its verification is worth. If nobody on your team can look at that output and say “that’s right” or “that’s wrong” with any authority, then you haven’t automated a task. You’ve just removed the last person who would have noticed the mistake.
That’s the ceiling. You cannot safely automate past the last person who can still grade the work by hand.
Here’s where it gets uncomfortable, and here’s why this is an owner’s problem rather than a staff problem.
A survey this year of 500 marketing professionals across five countries asked whether they’d ever acted on an AI recommendation they later suspected was wrong because the underlying data was bad.
Among individual contributors — the people closest to the actual work — 41% said yes.
Among C-suite executives, it was 78%.
Among senior vice presidents and VPs, 92%.
Read that again, because its shape matters more than the numbers. Exposure to incorrect AI answers doesn’t decrease as you move up the building. It more than doubles. The higher you sit, the more likely you are to have acted on something that turned out to be wrong.
And it isn’t because executives are careless. It’s structural. The further you are from the work, the less able you are to look at an output and feel that something’s off. The junior analyst who has pulled that report by hand 200 times gets an itch when a number looks strange. You don’t have that itch, because you were never supposed to need it.
Now stack a second study on top of that one. A separate executive benchmark, fielded this spring across 16 countries, asked leaders how confident they were in the accuracy of AI output that appears in critical external reports without a person reviewing it first.
84% were at least somewhat confident. 39% were very confident.
In the same study, only 11% agreed that their data quality was actually sufficient for AI use. And roughly 1 in 4 said their internal audits had already caught AI errors that reached outside audiences or reached the board.
So: 84% are confident in the output. 11% confident in the inputs. Two different surveys, two different populations, one finding.
Confidence is running years ahead of the ability to check.
The strategic implication here is not “slow down on AI.” It’s narrower than that, and more useful.
Your capacity to verify is a real business asset, and right now it’s almost certainly not on any list you keep. You know your headcount. You know your spend. You probably don’t know, department by department, where your grading capacity actually sits — who could reconstruct that output by hand if they had to, and what happens to that ability the longer the machine does it instead.
Because that’s the second-order problem nobody budgets for. Verification capacity isn’t fixed. It erodes. The skill that lets someone catch a bad answer is built by doing the work, and when the work goes to the tool, the skill goes with it. Automate the same step for eighteen months and the person who used to grade it can’t anymore. The ceiling drops while you’re standing under it.
And this connects directly to something we talk about constantly on this show. Proof over proximity. Producing evidence rather than claiming it. You cannot produce proof you are not able to verify. An unverified claim is not proof — it’s just a confident sentence, and confident sentences are the one thing these tools have in genuinely unlimited supply.
So here’s your self-audit for this week. Three ideas to think about:
First, name the last significant decision you made that began as an AI output. Then name the person who checked it before it reached you. If you can’t name that person, you’ve found something.
Second: Pick one process you’ve already automated. Ask whether anyone on your team could still do it by hand today — not whether they used to be able to, but whether they could today.
Third, and this is the one that actually predicts the next 12 months: when your team brings you the next thing they want to automate, notice whether anyone in the room asks who’s going to grade the output. If nobody asks, that’s not a tooling problem. That’s a standard you haven’t set yet.
A broken tool won’t tell you it’s broken. Someone has to be able to tell.
I want to thank you for listening. I’m Jeff Payne. I look forward to seeing you next time.
Your marketing team is quietly building software. The hours are coming out of the one asset that can’t be rebuilt later.
CONFIDENCE ISN’T A CHECK
A broken tool usually tells you it’s broken. That has been the arrangement between people and machines for roughly two centuries. Something fails, and something makes noise — a warning light, an error message, a line that stops moving. The failure announces itself,f and somebody goes to look.
AI does not participate in that arrangement. When an AI tool starts getting things wrong it does not slow down, flag anything, or sound less certain. It answers in the same tone it used when it was right: same speed, same clean formatting, same air of having considered the question carefully. You will not learn about the failure from the tool. You will learn about it from a customer, from a board member asking a follow-up nobody can answer, or — most often — not at all.
When an AI tool starts getting things wrong, it does not sound any less confident.
The constraint isn’t capability. It’s verification.
Nearly every AI conversation in a business right now is about scope. What should we automate, how much, how fast, and who owns it? Those are reasonable questions, but they arrive in the wrong order, because all of them assume the binding constraint is capability—what the tool is able to do.
It isn’t. The binding constraint is what your team can check.
Every AI output that enters your business arrives as a claim: a number, a recommendation, a draft, a summary. A claim is worth exactly what verifying it is worth. If nobody on your team can look at that output and say “that’s right” or “that’s wrong” with real authority, you have not automated a task. You have removed the last person who would have noticed the mistake.
That is the ceiling, and it is worth naming plainly: you cannot safely automate past the last person who can still grade the work by hand.
The exposure gets worse as you move up
This is where the data turns it from a staffing question into an owner’s question.
Validity’s State of CRM Data Report 2026 surveyed 500 B2B and B2C marketing professionals across the United States, United Kingdom, Brazil, Australia and New Zealand. It asked whether respondents had ever acted on an AI recommendation they later suspected was wrong because the underlying data was bad.
Among individual contributors — the people closest to the actual work — 41% said yes. Among C-suite executives, 78%. Among SVPs and VPs, 92%.
The shape of that finding matters more than the individual numbers. Exposure to wrong AI answers does not decrease as you move up the building. It more than doubles. And it is not a story about careless executives; it is structural. The further you sit from the work, the less able you are to look at an output and sense that something is off. An analyst who has pulled the same report by hand 200 times gets an itch when a number looks strange. Nobody above them has that itch, because nobody above them was ever supposed to need it.
Confident in the output, doubtful about the inputs
A second dataset, from an entirely different population, lands in the same place. Workiva’s 2026 mid-year executive benchmark — commissioned as an independent study through Ascend2 and fielded in May across 16 countries — asked executives how confident they were in the accuracy of AI output appearing in critical external reports without a person reviewing it first.
84% were at least somewhat confident, including 39% who were very confident. In the same study, only 11% agreed their data quality was sufficient for AI use, and roughly one in four said internal audits had already detected AI errors that reached external audiences or the board.
Set those two numbers beside each other.
84% confident in the output.
11% confident in the inputs.
Two surveys, two populations, one finding: confidence is running years ahead of the ability to check.
Confidence is running years ahead of the ability to check.
Verification capacity is an asset, and it erodes
The strategic implication is not “slow down on AI.” It is narrower and more useful than that.
Your capacity to verify is a real business asset, and it is almost certainly not on any list you keep. You know your headcount and your spend. You probably do not know, department by department, where your grading capacity actually sits — who could reconstruct a given output by hand if they had to, and what happens to that ability the longer the machine does the work instead.
That last part is the second-order problem nobody budgets for. Verification capacity is not fixed; it erodes. The skill that lets someone catch a bad answer is built by doing the work, and when the work moves to the tool, the skill follows it out.
Automate the same for 18 months, and the person who used to grant it can no longer do so. The ceiling drops while you are standing underneath it.
This connects directly to a principle this show returns to constantly: proof over proximity — producing evidence rather than claiming it. You cannot produce proof you are unable to verify. An unverified claim is not proof. It is a confident sentence, and confident sentences are the one thing these tools have in genuinely unlimited supply.
An unverified claim is not proof. It’s a confident sentence — and confident sentences are the one thing these tools have in unlimited supply.
Three questions to ask
First: Name the last significant decision you made that began as an AI output, then name the person who checked it before it reached you. If you cannot name that person, you have found something.
Second: Pick one process you have already automated and ask whether anyone on your team could still do it by hand today. Not whether they used to be able to. Whether they could today.
Third, and this is the one that actually predicts the next twelve months: when your team brings you the next thing they want to automate, notice whether anyone in the room asks who will grade the output. If nobody asks, that is not a tooling problem. That is a standard you have not set yet.
A broken tool will not tell you it is broken. Someone has to be able to tell.
Sources: This episode was prompted by the build-versus-buy analysis from Kevin Indig and Amanda Johnson at Growth Memo, “Are You Doing Marketing… or Building Software?” — specifically their rule that you should never buy or build a tool for a job nobody on your team can verify by hand. The seniority data comes from Validity’s State of CRM Data Report 2026. The executive confidence data comes from Workiva’s 2026 mid-year executive benchmark survey, conducted independently through Ascend2.
COMPLETE THE FORM TO
BOOK A STRATEGY CALL
"*" indicates required fields
COMPLETE THE FORM TO
BOOK A STRATEGY CALL
"*" indicates required fields
Subscribe and Share – WE APPRECIATE YOUR SUPPORT