Interviews
Kris Beevers, CEO and Co-Founder, Netbox Labs – Interview Series

Kris Beevers, CEO and Co-Founder, NetBox Labs is a technology entrepreneur and infrastructure software veteran with more than two decades of experience building companies and platforms focused on networking, cloud infrastructure, and automation. Before leading NetBox Labs, Beevers co-founded NS1 in 2013 and served as its CEO for nearly a decade, building the company into a prominent provider of network automation and application traffic management technology before its acquisition by IBM in 2023. As part of that transaction, NetBox Labs spun out of NS1 as an independent company with IBM as an investor. Earlier in his career, Beevers held senior engineering and architecture roles at Internap Network Services and Voxel, and he also co-founded SolidJoint Research.
NetBox Labs develops an infrastructure intelligence platform designed to help organizations model, operate, automate, and govern increasingly complex networks and IT infrastructure. The company is the commercial steward of NetBox, the widely adopted open-source network and infrastructure system of record used by more than 10,000 organizations. Its platform combines an infrastructure graph and source of truth with operational intelligence, automation, AI-assisted orchestration, and governance capabilities that allow both engineers and AI agents to safely interact with infrastructure. NetBox Labs supports cloud, self-managed enterprise, hybrid, and air-gapped deployments, while integrating with tools such as Ansible, Terraform, Nornir, and continuous integration and deployment pipelines.
You co-founded NS1 in 2013 and spent nearly a decade building the company before its acquisition by IBM, after which NetBox Labs emerged as an independent company. What lessons from building NS1 influenced why you founded NetBox Labs, and how has the infrastructure problem you are trying to solve changed in the AI era?
One thing I learned building NS1 is that infrastructure problems rarely stay neatly contained. DNS was our piece of the stack, but our customers were operating these incredibly complex environments where networks, data centers, applications and automation all depended on each other. The more time I spent with those teams, the clearer it became that understanding the infrastructure itself was a much bigger problem.
That was a big part of what drew me to NetBox. There was already this widely adopted open-source project and a community of engineers using it to model what they had, how it was connected and what it was supposed to look like. We saw an opportunity to build on that foundation.
What’s changed with AI is mostly the pace and scale. Infrastructure teams are being asked to build enormous environments incredibly quickly, while the underlying technology is changing just as fast. At the same time, we’re starting to automate more of the operation of that infrastructure which represents an exciting future. As AI is applied to infrastructure, IT teams are realizing they must have good, real-time data about their infrastructure to automate against, and they need to know what the intended state looks like so AI can help them identify when the operational infrastructure drifts from the plan.
So the lesson from NS1 still applies. Before you can automate infrastructure well, you have to understand it. AI just makes getting that right much more urgent.
For much of the last decade, cloud computing allowed developers and infrastructure teams to abstract away the physical hardware underneath their applications. Why is AI reversing that trend and forcing DevOps, Site Reliability Engineering (SRE), and network engineers to think about power, cooling, racks, cabling, and physical networking again?
Cloud taught a lot of us to treat infrastructure as effectively infinite. You asked for compute and it appeared. You didn’t necessarily have to care where the server was, how it was powered, how it was cooled or how all of the physical pieces underneath it came together.
AI infrastructure doesn’t really allow you to do that.
When you’re building these environments, you start with some pretty physical questions. How much land do I have? How much power can I get? What kind of cooling can I support? From there you get into racks, GPU servers, switches, fiber optic cabling and eventually the logical layer, IP addresses, configurations and software.
Those things all depend on one another. You can’t decide how many racks you’re going to deploy without understanding power and cooling density. You can’t think about the GPUs independently of the network connecting them.
That’s forcing disciplines that spent years moving away from the physical layer to engage with it again. The abstraction isn’t gone, but the physical constraints underneath it suddenly matter a lot more.
AI data centers are increasingly being discussed at gigawatt scale. What fundamentally changes operationally when infrastructure moves from conventional enterprise or cloud environments to facilities designed around enormous clusters of GPUs?
Gigawatt scale is absolutely astronomical. But while the scale is obviously different I think the more interesting difference is the amount of coordination required.
Think about what has to happen to bring a 300-megawatt data center online. You need land and power. Then you have to design the facility and procure racks, GPU servers, switches, fiber, power infrastructure and cooling equipment, often from completely different vendors with completely different ways of representing their products. All of that equipment has to arrive, get received, racked, cabled, configured, tested and ultimately handed over for training or inference.
And the ground is moving underneath you while you’re doing it. GPU architectures are changing. Networking is changing. Cooling requirements are changing. The components available six months from now may not be the same ones you designed around today.
So small inefficiencies compound very quickly. I recently spent time with one of the largest fiber optic cable manufacturers in the world, and they told me one of their biggest business problems is returns because customers order the wrong cable lengths. That sounds almost trivial until you’re ordering hundreds of thousands of cables.
At this scale, infrastructure operations become a giant logistics and constraint-satisfaction problem. The companies doing this well are the ones getting very good at carrying accurate design data all the way through procurement, deployment and operations.
You have said there is effectively no established playbook or talent pipeline for operating infrastructure at this scale. Which skills are currently hardest to find, and where do you expect the biggest talent shortages to emerge as AI infrastructure expands?
There are probably only a few hundred people in the world right now who really know how to build this kind of infrastructure at this speed and scale. And most of them are pretty busy actually doing it.
That’s part of what makes this moment unusual. There isn’t a mature body of knowledge you can just go study. The people doing this are learning from each other and figuring things out in real time. And because the technology is changing so quickly, some of those lessons become outdated pretty fast.
I think the shortage is therefore bigger than any one job title. We need people who understand networking, compute and automation, but increasingly also understand the physical environment those systems live in. Power, cooling, facility design, supply chain and field operations are becoming part of the same conversation.
The people who can cross some of those boundaries are going to be incredibly valuable. But I don’t think we’ve even settled on what all of those roles look like yet. The talent model is being built alongside the infrastructure.
As the boundaries between software, networking, facilities, energy, and data center engineering begin to blur, what new technical roles or hybrid skill sets do you expect to emerge?
I don’t think we know what all of those roles are going to look like yet. What we do know is that the people building this infrastructure have to think across a much wider set of problems than they did before.
You’re not just thinking about compute or networking in isolation. Power, cooling, physical design, supply chain, networking and automation all have to come together to get these environments online and keep them running.
I still think we’ll need people with deep expertise in each of those areas. But increasingly, they’ll also need to understand how decisions in their area affect the rest of the infrastructure. And because so much of this work has to happen faster, the ability to automate is going to matter across more of those disciplines.
AI agents are beginning to diagnose problems, generate configurations, and automate parts of infrastructure operations. Which responsibilities do you think AI will realistically take over from infrastructure engineers, and which will become even more dependent on deep human expertise?
I think a lot of the work where the inputs, desired outcome and boundaries are clear will increasingly be handled by AI. Generating configurations is an obvious example. So is diagnosing common problems, checking whether the infrastructure matches the intended design, or eventually fixing certain issues when there’s enough confidence about what went wrong and what the safe response is.
Where humans become more important is when the answer isn’t obvious.
Infrastructure fails in weird ways. A fiber gets cut. A device starts behaving differently than the design says it should. A change has an unexpected effect somewhere else in the environment. AI can help an engineer understand those situations much faster, but you still need people who understand the system deeply enough to decide what should happen next.
I think that’s the interesting shift. Engineers will probably spend less time doing repetitive configuration and troubleshooting and more time defining intent, designing systems, setting the boundaries for automation and handling the genuinely novel problems. All of that work will be augmented by AI, but driven by people.
That makes expertise more valuable, not less. The engineer who really understands why the infrastructure works the way it does is going to be incredibly important when the automation doesn’t have an obvious answer.
NetBox Labs has argued that AI systems managing infrastructure need an authoritative model of devices, connections, dependencies, and other physical and logical relationships. Why is this type of infrastructure context so important when moving from AI assistants that make recommendations to agents that can actually take actions?
The big difference is that once an agent can act, being wrong has real consequences.
An infrastructure agent needs more than a snapshot of what a device is doing right now. It needs to understand the environment around it: what exists, how things are connected, what changed recently and, importantly, what the infrastructure is supposed to look like.
Take something like troubleshooting a connectivity issue. It’s not enough to know that a device is unreachable. You want the agent to be able to trace the cable path, understand the dependencies around that device, look at recent changes and determine what else could be affected before it proposes what to do next.
That’s really the foundation we’ve spent years building toward at NetBox Labs, giving teams an accurate model of both the physical and logical infrastructure, along with the intent behind how it should operate.
But the data alone isn’t enough. You also need to decide what the agent is allowed to do on its own, what requires a person to approve it, and how every action gets tracked and validated.
Infrastructure isn’t like code where a bad change can always be cleanly reverted. A bad change can take down an operation. So as we move from AI that tells an engineer what it thinks to AI that can actually do the work, both context and control become much more important.
In your recent CIO article, “Why I, the CEO, am personally building our AI strategy,” you argued that AI is too important for company leaders to simply delegate and described personally prototyping with AI tools. How has being hands-on with these systems changed your thinking about what AI can realistically automate in infrastructure operations?
Being hands-on makes you much less interested in the theoretical conversation.
I’ve spent a lot of time actually building with these tools, most commonly these days prototyping or even building complete products with Claude Code. You learn pretty quickly that there’s a huge difference between seeing an impressive demo and building something you actually trust to do useful work.
You also develop a feel for where the technology is moving much faster than you can get from reading about it. Things I would have considered difficult to automate six months ago can suddenly be fairly straightforward. At the same time, you see very clearly where context, judgment and structure are still missing.
That’s influenced how I think about infrastructure operations. I’m very optimistic about how much operational work we can automate, but I think we’re distant from pure autonomy as the end goal.
The question I care about is much more basic. Does this help us operate infrastructure faster, more reliably or more effectively? If it does, great. If it doesn’t, it doesn’t matter how sophisticated the AI behind it is.
As AI data centers become increasingly constrained by electricity availability and cooling requirements, could infrastructure engineering evolve from primarily managing computing resources to actively coordinating workloads with energy and physical capacity?
Yes, and we’re already starting to see it. We have a phrase internally, “turbines in the parking lot,” that came from a real conversation with one of the teams building hyperscale AI infrastructure. They were bringing infrastructure online so quickly that the grid couldn’t keep pace, so they were literally buying turbines and putting them in the parking lot to get power fast enough.
That’s the kind of environment these teams are operating in. When power becomes one of your primary constraints, you have to get much smarter about how you use what’s available. Demand-side response, where operators actively coordinate AI workloads in reaction to shifts in the electrical grid, is already happening and is becoming an increasingly important capability for teams operating power-hungry infrastructure.
Not every workload has the same requirements. Latency-sensitive inference may need to stay online, while some training or batch workloads can be shifted or paused when power is constrained. I think we’ll increasingly see infrastructure teams managing compute, power and physical capacity as parts of the same operational problem.
Looking ahead, do you think the biggest bottleneck to scaling AI will ultimately be GPUs and model development, or the much broader challenge of finding enough power, physical infrastructure, networking capacity, automation, and skilled engineers to operate everything behind them?
I don’t think there’s going to be one bottleneck.
Building AI infrastructure at the speed and scale the market wants right now is basically a constraint-satisfaction problem. At any given moment, something is the primary constraint.
For a while, everyone was talking about GPUs. Power is obviously a huge bottleneck now. But it can also be networking equipment, cooling, land, fiber, procurement, construction, or simply finding enough people who know how to put all of this together.
And once you solve one constraint, another one becomes more visible. That’s what happens when demand is so much greater than supply.
So I wouldn’t bet on one permanent bottleneck. I think the more important capability is being able to adapt as the constraint moves.
That’s also why I don’t think anyone has the definitive playbook for AI infrastructure right now. The people building it are figuring this out while they’re scaling, and they’re doing both incredibly quickly.
Thank you for the great interview, readers who wish to learn more should visit NetBox Labs.












