The Funny Look Test: What to Ask When AI "Isn't Accurate Enough"


One p or two

“Philip with one l and one p,” my Dad said when we were filling in a form.

“Two ps,” I said.

“One p.”

“No, two. One at the beginning and one at the end.”

“Oh yeah.”

Dad was trained on a BSc and a PhD and he still hallucinates.

He laughed and accused me of being a smart arse.

Nobody walks away from that thinking Dad’s lost his marbles. Trust in a person was never “never wrong.” Trust in a tool shouldn’t be either.

Compared to what

Every time someone tells me AI isn’t accurate enough, I want to ask: compared to what?

Not compared to the expert you haven’t hired, haven’t got budget for, or couldn’t get an appointment with this quarter. Compared to the guess you’re making right now, on a spreadsheet, in a meeting, based on a hunch from three years ago that nobody’s revisited.

The numbers on human error aren’t flattering, and nobody applies the same scrutiny to them that they apply to AI:

Nobody’s proposing we ban spreadsheets. So “AI makes mistakes” can’t be the objection on its own. It’s not accurate enough compared to what.

The Funny Look Test

The question that matters isn’t how often it’s wrong. It’s whether it’s wrong about something material.

We know what good looks like in most of our own work. We know the order of magnitude a number should land in. Dave in Finance doesn’t need to be trusted instead of the AI. Dave needs to give the output “The Funny Look Test” before it goes anywhere near a decision. Does it look funny? That’s the whole test.

Back to Dad. He thinks he’s always right. He isn’t. But he’s right often enough to make you want to be on his team in a pub quiz. For the tasks at hand, his accuracy levels are good enough.

Silly Wild Ass Guesses

Whatever tool you use, the output is only as accurate as your riskiest assumption (AKA a SWAG = a silly wild ass guess). If that bad assumption is the thing that swings the decision, it doesn’t matter whether it came out of a slide rule, a spreadsheet, or an LLM writing Python. Most business problems are wicked, not tame: no single right answer, and the system changes the moment you act on it. What you need is a hypothesis about the future you’re willing to bet against, not a spreadsheet that makes you feel comfortable while it glosses over the assumption that matters most. That’s a value judgement, not a maths question. The expert stays in the loop.

What the objection is actually doing

A lot of the time, “not accurate enough” isn’t about accuracy. It’s deferring the decision. Or it’s someone exercising the only power they’ve got left in the room: the power to ask a question and have it factored in, whether or not the question is proportionate.

“We can’t possibly do this because customers in Outer Mongolia might be affected” sounds like rigour. It isn’t, when Outer Mongolia is 0.01% of revenue. It’s a stalling tactic dressed up as due diligence.

Ask it what you didn’t think to ask

At the scorecard workshop I wrote about last week, I asked the model what a malicious user might do to the tool I was building. The honeypot it suggested against bots isn’t something a naive user would have thought to build in. It didn’t hand me the comprehensive answer — nothing does. It closed part of a blind spot I didn’t know I had.

As I said at the time: it’s not perfect, but it’s better than us just sitting here kind of guessing.

Steal that move. Before you sign off on anything, ask it to name the risks you didn’t know to ask about. Then give the answer a funny look, and decide.

In a two-day hackathon, we don't just prototype random stuff, we take a step back and look at what needs to be in place to put this in production, including the funny look test, the blind-spot question, and your actual numbers and data quality, not someone else’s clean data set.

Till next time,

Helen

P.S. Two-day AI hackathon in September. Back-to-school energy, your team, working sessions not slideware. Book a call to talk more: https://calendly.com/helen-dawson/discovery-call

The Hard Part Newsletter

The Hard Part about adopting new digital tools and AI is almost never the technology, it's changing the way people work. This newsletter is for you if you're a leader struggling with where to start OR if some initiatives are running and you're wondering where the ROI is going to come from. Never more than a 5 minute read. Weekly.

Read more from The Hard Part Newsletter

Dear Reader, To quote the comedian, Vic Reeves: 88.2% of statistics are made up on the spot. In so many conversations that I have with leaders, they want numbers to give them confidence that they are on the right path with AI. Yet, CFOs operate on the basis that words talk, numbers scream. But is there any good data or is it all Lies, Damn Lies and AI Statistics, to paraphrase Disraeli (who was probably quoting someone else). How does this show up? Every month, reports claim "mind-blowing"...

I have a half-finished novel, a Creative Writing degree and have therefore read a lot about writing. This is known as procrasti-learning. So, borrowing from George Orwell's essay "Why I Write", I thought I'd chip in on the AI writing debate. Hence, my punny title. This comes in the week where Anthropic announced that they are "watermarking" the text that Claude generates. This is their response to the transparency requirements in the EU AI Act. Essentially, the watermark is using a non-random...

Dear Reader, "Where am I getting screwed?" that's the question entrepreneur Mark Cuban asks his large language model. I heard him say it in conversation with Allie Miller on a webinar this week. And it's the question behind one thing that really excites me from a consumer perspective about the potential of agentic AI: as a tool to combat the bureaucracy of every day life. No more negotiating labyrinth-like complaints processes, no more attempting to cancel Amazon Prime that you never really...