Thought Leaders
Your Systems Already Have Blind Spots. AI Just Makes Them Worse.

In 2022, before generative coding tools were part of our daily engineering work, I wrote about my philosophy on tool selection. It has held up better than I expected. Back then, I argued for starting with the problems you were actually solving, knowing your weaknesses, and prioritizing how you use the tools versus just jumping on whatever tool sounds best and hoping it works out. Know thyself, and thy goals, so you can set proper expectations for your tools.
At the time, I was thinking about SaaS sprawl, not AI-generated code. But today my philosophy is even more urgent, and more important to stand by.
Many of us have read the 2025 DORA report, which found that, unlike the year before, AI adoption now correlates positively with delivery throughput. The finding underneath it was that delivery instability kept climbing, and they tested whether the speed gains cover for it. They don’t. That matches our experience. Our team adopted agentic software development and saw a 48% increase in throughput over two quarters, followed by a 16% increase in stability issues. Ten people is a small sample, but it’s also a sample from which I can see the full picture, and the pattern held.
AI adoption isn’t really a question anymore. You’re either just starting out or you’re in the thick of it. What’s different now is that engineering leaders are expected to adopt AI and also prove that it’s paying off. The CEO, the board, and finance all want to know how to optimize their AI investment. They are asking whether the tools you picked are solving real problems efficiently.
The Gap Was Always There. AI Just Made It Wider.
As a CTO, I spend a good chunk of my time talking to other engineering leaders, including customers, prospects, and peers of mine to compare achievements and gripes about what we’re experiencing with AI. After enough of these conversations, I have started to see patterns in AI adoption and outcomes.
The main observation isn’t mine. DORA has led with it for two years now: AI amplifies whatever is already happening in the organization, both strengths and weaknesses. A team with clean architecture and healthy review habits gets faster. A team that’s wrangled a hairy ball of technical debt just enough to ship code now finds that tech debt is growing into a major blocker. What that framing misses, though, is why this catches so many teams off-guard. AI didn’t hide these weaknesses; the systems we’ve relied on never surfaced them.
The ticket-and-reporting stack most engineering orgs run on was built to answer questions that humans have, at human speed, by people who understood roughly what “done” meant for a given piece of work. It was never a perfect record. It was always an approximation, filled in by someone summarizing something messier underneath. Now, AI adds volume, and adds new inputs that generate new activity. None of the systems (or tools) used for traditional, non-AI methods of development were ever built for that.
Regardless, we are still responsible for the same objectives. You still own velocity, quality, spend, and how your team is actually doing. You just can’t take last year’s dashboards on faith anymore.
There’s a fair objection here. DORA’s 2026 ROI report describes a J-curve: a productivity dip right after adoption, driven by the learning curve, the cost of verifying AI-generated code, and downstream processes that haven’t caught up. They call it the “tuition cost” of transformation, and they warn leaders not to mistake it for failure. Fair enough. But tuition and a real problem look identical on a dashboard built from tickets. If you can’t tell which one you’re in, you’re not being patient. You’re guessing.
We need to go back to basics. Know thyself. Know thy team. Know what problems you are solving.
How Do You “Know Thyself” With AI?
From my conversations, I’ve identified five main areas where conventional systems, built for human-generated, human-reported work, are blind. Ignore them and you risk amplifying your weaknesses as you continue to adopt AI.
Blind Spot 1: Velocity Theater
More commits and more PRs can feel like progress, and often it is. AI raises both counts automatically. A Stanford case study showed that adopting AI lifted PR count 14%. But what you’re missing is how much of that activity is feature work that ships versus maintenance, rework, or churn from a refactor that didn’t stick.
To address this, watch the split between feature work and maintenance, and watch deployment frequency and lead time against your own historical baseline, not an industry average. Without that split, you’re reporting progress you can’t actually substantiate.
Blind Spot 2: Review Debt
Review capacity doesn’t automatically scale along with output. The verification tax isn’t a phase you get through; it’s part of the standing cost for agentic development. A recent survey of engineering leaders found 80% of teams spend at least 10% of their time on review, and roughly one in ten spend more than 40%. Under that load, teams swing between a growing backlog and rubber-stamping, and neither is a real answer.
The constraint on shipping isn’t how fast code gets written anymore. It’s how fast a human can actually be confident a change is correct, how fast and accurately defects can be detected and addressed. Watch how the review load is actually distributed across your team; otherwise you risk overloading your senior engineers, delaying your releases, or causing major production issues.
Blind Spot 3: Hidden Work
Refactors and architecture shifts have a habit of hiding inside other tickets, if they show up in the ticket system at all. AI produces more of this kind of work, not less. An agent doesn’t hesitate to touch twelve files to fix one bug, whereas a human might pause and reconsider. Work that skips the system of record also skips planning, which means your capacity model is wrong, and every forecast built on top of it is wrong too.
In order to understand how much work is actually being done, you need to watch how much is actually changing in the codebase and in pull request history. Without that, your capacity plan is built on what people remembered to log, not on what they actually did.
Blind Spot 4: Quality Drift
The same survey found nearly half of engineering leaders struggle to detect security issues week to week. Complexity, duplication, and dependencies that don’t quite belong build up across a lot of small, individually reasonable changes. None of them look alarming on their own. In the same Stanford case study, code quality dropped 9% and its variance more than tripled. While the average moved a little, the spread (the part you notice) moved a lot. At AI volume, they compound faster than most review processes catch them. Drift tends to surface as an on-call alert traced back to a dependency nobody remembers reviewing. By the time this happens, there’s a decent chance a customer has noticed first.
Watch the trend lines on security findings, dependencies, and failure and recovery—not the individual commit. Complexity and duplication creeping up over several weeks matter more than any one change that gets flagged in review. Without that, you catch drift the way most teams still do: after it’s already caused an incident.
Blind Spot 5: Unproven Spend
Once AI adoption stops being a debate, AI spend and ROI become the question everyone focuses on. Finance wants to know what’s capitalizable versus operational. Leadership wants to know what the investment resulted in. Most teams are still making tooling, seat, and headcount decisions based on intuition, not proof of the link between money and delivered work.
Watch where engineering effort actually flows in the codebase itself, quarter over quarter — not where the roadmap says it’s supposed to flow. Without that link, you’re defending next year’s budget with anecdotes, and anecdotes don’t survive a hard conversation with a CFO.
Start With What You Can’t See
Finance’s question about capitalized spend and an on-call page at 2 a.m. look unrelated, but aren’t. Both can be “guesstimated” from activity. But both are actually answerable, with evidence, from the code itself.
DORA’s answer to all of this is the engineering system itself: platform quality, workflow clarity, team alignment. That’s right, and it’s also not step one. You can’t fix a system you can’t see. Every one of those five areas is something you have to be able to observe before you can argue for investing in it.
The first useful step isn’t a new tool or a new process. It’s knowing thyself, honestly, and determining which of these five areas are blind spots where you lack real evidence. Most leaders can immediately identify (and are paying attention to) problems in one of these areas. However, it’s the areas where you have the least information that are most likely to come up and bite you as you continue to adopt AI.












