
What’s Changing in AI Infrastructure | ALL IN 2026
Explore the AI infrastructure themes that stood out at ALL IN 2026, from GPU capacity and Kubernetes to data residency, operations, and production readiness.
Read field noteField notes / Latest
Engineering notes from operating open infrastructure: the failures, design decisions, and upstream work that make open infrastructure better.
Browse all field notes
Explore the AI infrastructure themes that stood out at ALL IN 2026, from GPU capacity and Kubernetes to data residency, operations, and production readiness.
Read field note
VEXXHOST reflects on ALL IN 2026 in Montréal, sharing conversations around AI infrastructure, GPUs, Kubernetes, data residency, and production readiness.
Read field note
What makes private cloud difficult to operate? Explore Day 2 operations, upgrades, automation, capacity, skills, and managed cloud options.
Read field noteExplore the AI infrastructure themes that stood out at ALL IN 2026, from GPU capacity and Kubernetes to data residency, operations, and production readiness.
ALL IN 2026 brought together people building, deploying, and supporting AI across a wide range of organizations. For VEXXHOST, some of the most valuable moments came through conversations with teams working through real infrastructure questions, from GPU access and Kubernetes to storage, networking, data location, and operations.
One theme came through clearly: AI is increasingly becoming an infrastructure challenge.
The conversation is shifting from whether organizations can run AI to whether they can operate it reliably, efficiently, and at scale.
Here are five infrastructure insights that stood out at ALL IN 2026.
GPUs understandably receive a lot of attention. But having access to accelerators does not automatically create a production ready AI environment.
AI workloads depend on the systems around those GPUs: networking, storage, orchestration, scheduling, monitoring, security, and data movement.
A team can add more compute and still encounter poor utilization or performance if data cannot reach workloads efficiently, resources are difficult to schedule, or infrastructure teams lack visibility into what is happening across the environment.
As AI adoption grows, the conversation needs to expand from GPU capacity to infrastructure capacity.
The question is not simply how many GPUs are available. It is whether the entire infrastructure stack can support the workloads expected to run on them.
For VEXXHOST, that means looking at AI infrastructure as a complete environment, not a single component. Compute, networking, storage, Kubernetes, and managed operations all need to work together if teams want to scale with confidence.
Another valuable insight is that simply adding capacity does not remove operational complexity.
When a handful of experiments become shared environments used by multiple teams, new questions appear quickly.
How should GPU resources be allocated? Which workloads require dedicated resources? What happens when demand exceeds capacity? How are costs measured? How are infrastructure changes tested before they reach production?
These decisions are easier to make before an environment grows.
Teams preparing to scale AI infrastructure should establish clear approaches to capacity planning, workload placement, observability, lifecycle management, access, and recovery early.
Doing that groundwork creates a much stronger foundation than trying to introduce operational discipline after infrastructure has already become difficult to manage.
As organizations move toward shared and production AI environments, Kubernetes continues to appear naturally in infrastructure discussions.
It can help standardize deployments, schedule workloads, isolate applications, automate recovery, and provide a consistent operating model across environments.
But Kubernetes should not be viewed as the entire AI platform.
Its effectiveness still depends on what sits underneath it: compute, GPU integration, networking, storage, monitoring, and the processes used to operate and upgrade the environment.
The goal is not simply to say that AI workloads run on Kubernetes.
The goal is to create a stack where the infrastructure and orchestration layers work together predictably, particularly as more teams and applications compete for shared resources.
That is where the broader infrastructure foundation matters. Kubernetes becomes much more effective when it is supported by reliable compute, networking, storage, and operational processes underneath it.
Performance is only one factor shaping AI infrastructure decisions.
Organizations are also thinking about where workloads and data live, how portable their environments are, and how dependent they want to become on a particular infrastructure model.
For Canadian organizations, data residency and infrastructure location can be important parts of that discussion. At the same time, teams need flexibility because the AI ecosystem itself is changing quickly.
The models, accelerators, tools, and deployment patterns organizations use today may not be the ones they rely on several years from now.
That makes infrastructure control and portability valuable architectural considerations.
Open technologies such as OpenStack, Kubernetes, and Ceph can provide building blocks for environments designed around flexibility, portability, and long term control rather than unnecessary platform dependence.
For organizations that want to keep workloads and data in Canada, these considerations can become even more important as AI moves closer to business critical systems.
AI pilots are usually built around one primary goal: prove that something works.
Production infrastructure has a much longer list of requirements.
Can workloads recover when something fails? Can teams understand where GPU capacity is being consumed? Can infrastructure be upgraded without unexpectedly breaking applications? Can engineers identify networking or storage bottlenecks? Can the environment scale without costs becoming difficult to understand?
These are day two operational problems.
They may receive less attention than new models or accelerators, but they become increasingly important as AI workloads move closer to critical business systems.
For infrastructure teams, that means observability, lifecycle management, upgrade planning, workload isolation, capacity management, and operational consistency need to be considered part of AI readiness.
This is also where managed operations can make a significant difference. As environments become more complex, teams need not only the right infrastructure, but also the processes and expertise to operate it reliably over time.
Our biggest takeaway from ALL IN 2026 is that the AI infrastructure conversation is becoming more mature.
The question is moving from:
“Can we run AI?”
to:
“Can we run it reliably, efficiently, and at scale?”
Answering that question requires looking beyond individual GPUs or individual applications.
Compute, networking, storage, Kubernetes, observability, lifecycle management, cost visibility, data location, and infrastructure control increasingly need to be treated as parts of the same system.
For organizations moving from isolated AI experiments toward shared or production environments, this is a good moment to evaluate the foundation underneath those workloads.
The opportunity is to build that foundation deliberately, with the flexibility, visibility, and operational control needed to support what comes next.
As AI becomes more important to the business, the infrastructure underneath it becomes more important too.
Thanks to everyone who stopped by the VEXXHOST booth during ALL IN 2026 shared what you are working on, and gave us a closer look at the infrastructure questions teams are trying to solve.
Choose from Atmosphere Cloud, Hosted, or On-Premise.
Simplify your cloud operations with our intuitive dashboard.
Run it yourself, tap our expert support, or opt for full remote operations.
Leverage Terraform, Ansible or APIs directly powered by OpenStack & Kubernetes