What the research keeps confirming.
A client of mine has a coaching agent, a voice model, and a full content system already built. She still can't get away from her desk.
Every piece of content that goes out still routes through her twice. Once to get assigned. Once to get rewritten, because whoever wrote the first draft couldn't sound like her. She has the tools. She built the tools with me. And she is exactly as bottlenecked as she was before any of it existed, just with more software open.
I sat with that for a while before I said anything to her. Because my first instinct was to think something was wrong with the tools. It wasn't. The tools were fine. What was underneath them wasn't.
I went looking for whether this was just her, or whether it was a pattern. It's not just her.
MIT's Project NANDA found that 95% of generative AI pilots inside companies produce no measurable return, even with $30 to $40 billion in enterprise spend behind them. Not a modeling problem. Their own language for it is a learning gap: the tools don't retain what they're taught, they don't adapt from one interaction to the next, and most of all, they don't integrate into how the business actually runs. One manufacturing executive they interviewed put it plainly: the hype on LinkedIn says everything has changed, but in his operations, nothing fundamental has shifted.
PwC surveyed over four thousand CEOs and got a similar number from a different angle. Fifty-six percent saw zero cost or revenue improvement from AI. The pattern PwC named was deploy first, ask what problem you're solving second.
Terminal X went a layer deeper, and this is the piece that actually reframed it for me: a way to measure whether anything worked, the infrastructure for the tool to actually connect to real workflows, and a strategy that treats the rollout as something the whole org has to metabolize, not just install. Skip any of those three and it doesn't matter how good the model is.
And the most recent number, from a McKinsey survey published two days ago as I'm writing this: only 37% of companies can attribute any bottom-line impact to AI at all, flat since last year. Only 6% count as real high performers. McKinsey's explanation for the gap between that 6% and everyone else wasn't tool quality. It was organizational change. The 6% redesigned how decisions get made and who owns what. Everyone else added a tool to the org chart they already had.
That's six independent reports, four different research groups, two different years, all landing on some version of the same sentence: the system it's landing on determines whether AI does anything at all.
Two ways a system can be broken.
Once I had language for this, I started checking it against actual clients instead of just my one example. And what I found is that "broken system" isn't one thing. It's at least two, and they need completely different fixes.
Shape one: the bottleneck
This is my client above. Everything has to pass through one person because nothing downstream can act without her. There's no single place a decision lives once it's made, so it lives in whoever remembers making it. Nobody besides her has standing to move a piece of work forward without her sign-off. And "sounding like her" only exists in her head, so every AI output is still a first draft that has to come back to her to become real. Add AI to that shape and you haven't removed the bottleneck. You've given it more to check.
Shape two: the blind spot
I saw this with a different client, a small coaching-certification business. Their website broke four separate times over several months, including an integration that silently rerouted every new lead to the founder personally instead of the round-robin it was supposed to hit. There was a named owner for the website. He'd taken full responsibility for it. Ownership wasn't the problem. Detection was. Nobody found out for months, because nothing was checking whether the funnel actually worked.
The founder discovered it the way most people discover a blind spot: she noticed the effect, a sudden cash crunch, months after the actual cause. AI can't fix a hole nobody can see, and neither can a new hire, or a better tool, or more spend.
Both of these get called "our systems are broken." Both make new AI investment a waste until the specific missing piece gets installed. But they're not the same missing piece.
The five questions I actually ask now.
Before I recommend any AI tool to a client, I check for which shape they're in first. Sometimes it's both.
- Is there one place a decision lives once it's made, or does it live in whoever remembers it?
- Does anyone besides the founder have standing to move work forward without her sign-off on every draft?
- Does the standard, the thing that makes it sound and look like them, exist as something an AI tool could actually be pointed at, or does it only exist in their head?
- Is there a baseline and a check that would surface a failure within days, or does it only surface once the damage has already piled up?
- If something is working, can you point to the number that proves it, or is "it seems fine" the whole measurement?
None of these questions are about AI. That's sort of the point. AI amplifies whatever is already true about how the business runs. Layer it onto a bottleneck and the founder becomes the bottleneck faster. Layer it onto a blind spot and it fails the exact same silent way the last three tools already did.
The part I haven't fully resolved.
Here's something I'm still sitting with, and I'd rather say it out loud than pretend I've got it figured out.
Clients almost always want to hire before the system underneath is solid. A social media person, an ops hire, someone to take something off their plate. My instinct is to tell them to pump the brakes and fix the system first. But that instinct makes me nervous too, because it puts me in the position of being the one holding everything until it's ready, instead of them just going and getting the help. I don't fully trust that a new hire walks into a broken system and doesn't just become part of the breakage. I also know I could train one.
That's not overcaution. That's the single most correlated factor in whether a company survives its own growth.
There's research that backs the brakes instinct even outside of AI specifically. Startup Genome studied over three thousand high-growth companies and found that 74% of startup failures trace back to premature scaling, meaning the company invested in the next stage, more spend, more hires, more surface area, before the current stage was actually solid. The ones that sequenced it correctly grew roughly twenty times faster than the ones that didn't.
So maybe the reframe is: I'm not competing with the future hire. I'm the reason the hire succeeds instead of inheriting a mess. I'm not fully settled on that yet. I wanted to include it anyway, because it's the same pattern from the inside, and I think naming the tension is more honest than pretending I arrived somewhere clean.
Where this actually goes.
If you're staring at an AI tool that isn't doing what you hoped, the question probably isn't which tool to try next. It's which of these two shapes your business is actually in, and whether anyone has looked.
I've started calling this a systems check. It's really just an hour where I walk through the five questions above against how your business actually runs today, not how you'd describe it on a good day. Most people leave it with a clear answer to something they'd been vaguely worried about for months.
The tools were never the problem. The question is just whether anyone checked what they were landing on.