Thought Leaders
The Hidden Cost of AI: How “Black Box” Models Are Eroding Trust, Budgets, and the Environment

The AI industry has a model design problem. Enterprises are starting to feel it in cost, performance, and reliability.
AI can talk – but can it listen?
Listening is far more than simply processing words; it’s about interpreting signals and understanding nuance to make decisions in real time. But most “voice” models are built to transcribe tokens, not to recognize the subtle meaning hidden in a conversation. That means they’re missing crucial information.
In environments like voice, context, timing, and continuous input matter. As AI systems move from demos to real-world use, a deeper issue in how these systems are built is becoming clear, and enterprises are starting to feel the impact of those gaps across cost, performance, and reliability.
The dominant approach in AI today prioritizes larger, more general-purpose models. But as those models are deployed at scale, their limitations become harder to ignore. They’re expensive to run continuously, difficult to interpret and trust, and often inefficient for the tasks they’re applied to. What looks like an incredible one-off demo can break down when deployed at enterprise scale, where systems need to be responsive, efficient, and understandable all at once.
This dynamic creates a tradeoff: more optionality in what the model can do in theory, but at the expense of efficiency, transparency, and ultimately, usability. The cost of running massive models at scale is climbing, the energy footprint is (justifiably) under scrutiny, and the systems themselves are becoming harder to interpret. Before we reach the point of no return, we should ask ourselves: is this really the most sustainable way to build AI systems?
Why General-Purpose AI Models Are Inherently Inefficient
The industry’s push toward general-purpose AI has created a new kind of inefficiency: using the same heavyweight model for every problem or task, regardless of complexity. It’s an appealing idea – a single system that can handle anything – but in practice, it’s a blunt instrument.
Not every task requires a massive model. In fact, most don’t. When organizations rely on monolithic systems, they end up paying the cost of that complexity on every interaction. Simple decisions are processed with the same computational overhead as complex ones, and the economics quickly stop making sense at scale.
This is especially true in real-time environments like voice, where systems need to process continuous streams of data. Applying a large, general-purpose model to every moment of every interaction is both expensive and unnecessary. The more broadly these models are deployed, the more that inefficiency adds up. We end up with the worst of both worlds: systems that are too expensive to run everywhere, and too limited to be effective where they are.
A clear example of this shows up in contact centers. Many organizations would like to apply AI to every customer interaction, but the cost of running large models continuously makes that impractical. As McKinsey notes in its Global Survey on AI, organizations continue to struggle to scale AI use cases broadly due to cost and operational complexity (to put it into perspective, large enterprise contact centers often process 5,000 to 15,000 customer calls per day, adding up to millions or even billions of interactions each year). To manage this volume, they analyze only a subset of calls or rely on post-incident review. The downstream impact is measurable: 39% of organizations report higher call volumes tied to fraud-related inquiries, 34% report longer handling times, and 29% report reduced employee productivity¹. The result is a system that’s both overpowered and underutilized: expensive to run yet unable to deliver full coverage or real-time insight.
This is where overgeneralization becomes a liability. When a system is designed to do everything, it often ends up doing each thing inefficiently. The result is wasted compute at scale, running expensive models where simpler, more targeted approaches would be more effective. As organizations push to apply AI more broadly, that inefficiency compounds quickly, making it harder to justify continuous, real-time use.
The Black Box Problem: When AI Can’t Explain Itself
As models have grown more complex, they’ve also become harder to interpret. In many cases, AI systems can deliver mostly accurate outputs, but provide little insight into how those decisions were made, or what caused the remaining errors. That obfuscation slows down the people responsible for acting on the model’s results, losing the advantage gained from the speedup in the first place.
This is especially true in domains such as fraud detection, where decisions must be made in real time. If a system flags a call as high-risk but can’t explain why, it creates hesitation at exactly the moment speed matters most. Teams are forced to choose between trusting a system they don’t fully understand or second-guessing it in real time. Both options introduce risk. As AI moves deeper into operational workflows, the ability to understand and respond to model outputs in real time becomes a core requirement, not a nice-to-have.
This is where the limitations of large, general-purpose models become more apparent. Their complexity makes them harder to interrogate, harder to debug, and harder to adapt. As a result, organizations aren’t just dealing with higher costs; they’re dealing with systems that are less aligned with how decisions actually get made.
Where This Breaks Down: Real-Time AI in the Real World
These challenges – cost, inefficiency, lack of transparency – are manageable in isolated use cases. They become much harder to ignore in real-time systems, where AI isn’t just generating outputs, but actively shaping decisions as they happen.
Voice is a good example. Unlike batch processing or offline analysis, voice interactions are continuous and time-sensitive. Systems don’t get to pause, reprocess, or escalate decisions later; they have to operate in the moment. An AI system misunderstanding “Oh good, another AI” as positive when it was spoken sarcastically may be the difference between keeping and losing a customer, and must be acted on in seconds, before any human can intervene. That means running models continuously across entire conversations while also producing outputs that humans can act on immediately.
This is where the current paradigm starts to break. Large, general-purpose models are expensive to run at that scale, and their outputs are often too opaque to support real-time decision-making. As a result, companies are forced to make compromises: analyzing only a subset of interactions, delaying decisions until after the fact, or relying on imperfect human judgment in situations where speed and accuracy are critical.
In other words, the limitations of today’s models aren’t theoretical – they show up most clearly in the environments where AI is expected to deliver the most value.
A Different Path Forward: Smarter – Not Bigger – AI
If the current approach to AI is defined by scale, the next phase will be defined by efficiency.
Instead of routing every task through a single model, systems can break workflows into stages, applying lightweight models for routine decisions and reserving heavier processing for complex cases.
This subtle shift in architecture changes the economics. When systems can dynamically decide how much processing is required for a given task, they avoid the overhead of running large models unnecessarily. That makes it more practical to apply AI continuously, rather than selectively. It also creates opportunities for greater transparency, since decisions can be tied to specific signals or components within the system, rather than emerging from a single opaque model.
For enterprises, this is what makes broader adoption possible: The ability to analyze more data in real time without prohibitive cost or complexity unlocks use cases that have been out of reach.
The industry has made significant progress by growing models – but one size doesn’t fit all. As AI moves deeper into operational systems, the question is shifting from what models can do to how they should be designed. The future of AI won’t be defined by the largest models, but by the systems that can deliver intelligence efficiently, transparently, and at scale.












