Thought Leaders
Stop Designing AI Infrastructure Around the GPU

Why MSPs Should Start with the Workload, Not the Hardware
Spend five minutes at an AI conference and you could easily come away believing that every successful AI deployment begins with buying more GPUs. It is easy to understand why. Hardware dominates the conversation. Customers hear about Blackwell systems, InfiniBand fabrics, hyperscale clouds, and increasingly massive AI clusters. Vendors naturally gravitate toward the newest accelerators and fastest systems because they are exciting, relevant, and relatively easy to place in the market.
The problem isn’t that compute doesn’t matter. It matters enormously.
The problem is that starting there can lead organizations to ask the wrong question. The AI market is no longer in the experimentation phase. AI is getting put into production, companies are investing real money, and they’re expecting measurable business outcomes. Infrastructure decisions have become a great deal more consequential than they were even two years ago. Yet, not enough decisions are being driven by business requirements – technology led decisions remain in the lead.
The first question should not be “Which GPU should we buy?”
“What workload are we trying to support?” should be the focus.
That seemingly small change affects almost every infrastructure decision that follows.
There Is No Standard AI Infrastructure
One of the biggest misconceptions in the market is that there is a standard blueprint for AI infrastructure. There isn’t.
We talk about AI as though it were a single workload. In reality, AI encompasses an enormous range of business applications with very different requirements. A voice AI platform doesn’t have the same infrastructure requirements as medical imaging. Knowledge retrieval is different from image generation. Fraud detection doesn’t look like predictive analytics, and neither resembles video processing. They all use AI. They simply use infrastructure differently.
You’re not really designing infrastructure for “AI.” You’re designing infrastructure for a business application that happens to use AI. That distinction matters. Each and every workload puts unique demands on the infrastructure supporting it. Some require substantial compute resources. Others depend heavily on storage performance because they’re continuously retrieving large datasets. Some are constrained by network throughput, while another might live or die on latency because every millisecond affects the customer experience.
There is also a practical reality. The infrastructure a model was designed for isn’t always the infrastructure available when it’s time to deploy. Hardware availability, long lead times, or deployment deadlines can force organizations to use different GPUs, accelerators, or infrastructure configurations than originally planned. That can mean re-optimizing the model. Or even redesigning the model around the hardware they can actually deploy.
Security and governance requirements are equally workload-specific. An application processing public information has very different requirements from one handling financial transactions, healthcare records, or proprietary intellectual property. Data protection, identity and access management, compliance, sovereignty, backup, recovery, and availability cannot simply be added after deployment. They are architectural decisions.
Business requirements add another layer. How quickly will the application need to scale? What operating costs are sustainable? What availability level does the business require? How much complexity can the organization realistically manage? Those questions will be answered differently by each customer. That’s why there is no one-size-fits-all AI infrastructure.
Organizations beginning with a preferred cloud, hardware platform, or vendor aren’t getting AI infrastructure right. The leaders are beginning with the workload and designing an architecture around the business objective.
Training Gets the Headlines. Inference Delivers the Business Value.
The industry’s fascination with training is another reason AI infrastructure conversations can head in the wrong direction.
Training a large language model is an extraordinary engineering challenge. Enormous datasets, massive GPU clusters, significant power, and infrastructure capable of operating at full capacity for days, weeks, or even months are required. It is expensive, technically impressive, and naturally attracts attention.
Most organizations, however, aren’t building the next frontier model. They’re building customer service applications, voice AI systems, employee copilots, knowledge assistants, search tools, document summarization platforms, fraud detection systems, and dozens of other practical applications using models that have already been trained.
Those are inference workloads, and inference changes the infrastructure equation. Instead of optimizing exclusively for maximum compute, organizations may need to optimize for fast response times, low latency, predictable operating costs, and consistent performance.
A customer doesn’t care how powerful the underlying GPU is if a chatbot takes five seconds to respond. A caller doesn’t care about the specifications of the AI cluster if a voice assistant repeatedly misunderstands requests or hesitates during a conversation. They simply know the application isn’t performing well.
Designing every AI environment as though you’re training a foundation model is therefore usually the wrong approach and frequently an unnecessarily expensive one.
The goal of most MSP customers is not to build the world’s largest GPU cluster. Getting AI applications into production quickly, reliably, securely, and economically is the goal.
The challenge is finding the right balance of performance, security, scalability, resilience, and cost for the workloads they’re actually running.
Maybe the GPU Isn’t Your Bottleneck
GPUs have become the celebrity of AI infrastructure. They’re expensive, difficult to obtain, and easy to compare, which makes them the centerpiece of countless infrastructure conversations. The GPU may not be the thing holding it back once an AI application reaches production, however.
“How many GPUs do we need?” is not the question we should be asking, rather it is “What’s going to slow this application down six months from now?”
The answer may also be somewhere else in the architecture.
Storage is a good example. Enormous amounts of data are consumed by AI workloads – and, those datasets grow over time. Even an extremely powerful GPU can spend valuable time waiting rather than working, if storage can’t deliver information quickly enough. That data also needs to be protected, backed up, retained, secured, and managed throughout its lifecycle.
Just as much, networking matters. Throughput, latency, east-west traffic, and communication between AI clusters all affect application performance. A well-designed compute environment cannot compensate indefinitely for a poorly designed network.
Also, security has to be part of the architecture from the beginning. Questions that must be addressed prior to production include: where sensitive data resides, how networks are segmented, whether workloads communicate over private or public connectivity, and how compliance and sovereignty requirements are addressed.
Another easily overlooked factor is connectivity. While they may not generate flashy headlines, fiber diversity, route diversity, peering relationships, and geographic proximity, can critically influence the user experience – not to mention platform resilience.
End customers don’t know and don’t care which GPU is sitting in the rack. They care whether the application responds immediately or leaves them waiting.
Physical infrastructure deserves attention too. Power availability, cooling capacity, rack density, and expansion capacity determine whether today’s successful deployment can accommodate tomorrow’s growth.
Then there is data gravity. As datasets expand, moving petabytes of information between locations, simply because the compute happens to reside somewhere else, becomes increasingly inefficient. In many situations, bringing compute closer to the data can be both more practical and less expensive.
This is why architecture matters.
Think about a race car – just because it has the best engine doesn’t mean it is going to win. The transmission, tires, suspension, track, and especially the driver all matter too. AI infrastructure works much the same way.
The organizations generating the greatest value from AI won’t necessarily be those with the biggest GPU clusters. They’ll be the ones that understand how every layer of the infrastructure works together.
That’s the difference between buying infrastructure and designing it.
A Workload-First Planning Framework
MSPs have an opportunity to change the infrastructure conversation.
Instead of beginning with:
- Which GPU?
- Which cloud?
- Which vendor?
Start with the workload:
- What business problem are we solving?
- Is this a training or inference workload?
- How much latency can the application tolerate?
- Where does the data live, and how quickly will it grow?
- What security, compliance, and sovereignty requirements apply?
- How will the workload scale?
- What level of availability does the business require?
- What level of operational risk is acceptable?
- What will this environment cost to operate as usage grows?
The answers should determine the architecture. Not the other way around.
The Opportunity for MSPs
This shift changes the role of the MSP.
Customers don’t need another partner capable of selling them infrastructure. A partner capable of helping them make better infrastructure decisions is what’s needed.
A workload-first approach is a must as it gives MSPs the opportunity to evaluate compute, storage, networking, connectivity, security, data location, availability, and cost as parts of a single architecture – rather than separate purchasing decisions.
In this way, you’re able to control costs, improve performance, and identify operational and security risks before applications reach production.
A better business model for the MSP is also created.
MSPs can build higher-value recurring services around architecture, deployment, optimization, security, lifecycle management, capacity planning, and continuous improvement – instead of competing primarily on shrinking hardware margins.
The value isn’t in recommending the latest GPU or newest cloud platform. It’s in knowing when a customer needs them, when they don’t, and what else must be designed around them.
AI infrastructure is ultimately not a hardware decision. It’s an architecture decision driven by the workload, the data, and the business outcome the customer is trying to achieve.
The MSPs that understand that distinction will be positioned to become something much more valuable than infrastructure suppliers.
They’ll become the people customers trust to help decide what infrastructure they actually need.












