Thought Leaders
AI Isn’t Ready for the Real Work: Why Models Flunk Complex Tasks

Investors in the AI bubble beg white collar professionals to hand over their hardest problems and promise workers that they’ll get their afternoons back if they just give up a little more of their data. Complex analysis, judgment calls, and the workflows that took years to develop are absorbed by a model, supposedly freeing humans for something better. This is the dominant narrative in corporate America right now, and it is a trap.
The Capability Myth: Why Models Flunk Complex Tasks
A new study out of UC Berkeley tested leading AI agents against real world professional tasks across 55 industries, from healthcare to finance to law. Overall performance came in below 25 percent. Then researchers isolated the hardest task in each field – the kind of work that separates a senior professional from a junior one – and every single model failed. The scores weren’t just low. They were at zero.
You would think that a result like that would have changed the narrative on its own, but it hasn’t. Across industries, AI investors have too much riding on convincing professionals that most current AI models are further along than they are. But the most interesting question isn’t whether AI can do the hardest parts of a job today, because we, and everyone selling AI, already know the answer to that.
So, this leads to another question. What happens to a young professional’s judgment when they buy into this data exchange, before their own reasoning has fully formed?
The Minnesota Findings: How AI Flattens Skill Toward the Middle
A new empirical study out of the University of Minnesota Law School offers a real answer, using law as the test case. Researchers ran about a hundred upper-level law students through a sequence of realistic tasks: synthesizing law from raw source material, answering closed book questions to test what they absorbed, applying that law to a fact pattern, and then revising their own work. One group had AI access only at the very first and very last steps. The other had no AI until the end.
Students who used AI to build their first draft produced stronger work, and produced it faster, which isn’t entirely remarkable. What is remarkable, however, is what came after. The researchers expected early AI use to leave a comprehension gap that would show up later when the tool was taken away. But this didn’t happen. Students who had used AI to synthesize the law actually understood it better than the students who worked without it, once tested cold. Independent reasoning didn’t erode. It held, and in some cases it improved.
The twist came at the final stage, when every student had AI available to revise their work. Students who started with a weaker memo used that access to genuinely improve it. Students who started with a stronger memo got worse. Left alone with a capable tool and no more structure around how to use it, skill didn’t compound but flattened toward the middle.
These two studies together provide a stark truth: AI is not a replacement for professional judgment, and the Berkeley numbers should end that conversation for good. But AI is also not some corrosive force that eats away at reasoning on contact. Its effect depends entirely on where it shows up in the process.
Sequencing Over Substitution: Keeping Judgment at the Core
This applies outside the practice of law. Every knowledge profession is running its own version of this experiment right now: the doctor deciding whether to let a model draft a differential diagnosis, the analyst deciding whether to let a model build the first pass of a model, and the associate deciding whether to let a model structure an argument before they have worked through it themselves. The Minnesota results suggest the answer is not a blanket yes or no to all forms of AI. It is a question of sequencing. AI belongs on the front end of the grunt work and the parts of a job that generate friction, but not judgment. It does not belong at the center of the task the client is actually paying for, and right now it cannot get there even if it tried.
The companies still promising otherwise are selling a version of AI that the data says does not exist. The ones who build their products, and their client relationships, around the line the research actually draws will be the ones professionals still trust in five years. Everyone else is going to spend that time explaining a very expensive misunderstanding.












