
OpenStack for AI: How to Build a Private GPU Cloud for Production
Learn how OpenStack can power a private GPU cloud for enterprise AI, including compute, networking, storage, Kubernetes, isolation, operations, and deployment choices.
Read field noteField notes / Latest
Engineering notes from operating open infrastructure: the failures, design decisions, and upstream work that make open infrastructure better.
Browse all field notes
Learn how OpenStack can power a private GPU cloud for enterprise AI, including compute, networking, storage, Kubernetes, isolation, operations, and deployment choices.
Read field note
Explore the AI infrastructure themes that stood out at ALL IN 2026, from GPU capacity and Kubernetes to data residency, operations, and production readiness.
Read field note
VEXXHOST reflects on ALL IN 2026 in Montréal, sharing conversations around AI infrastructure, GPUs, Kubernetes, data residency, and production readiness.
Read field noteLearn how OpenStack can power a private GPU cloud for enterprise AI, including compute, networking, storage, Kubernetes, isolation, operations, and deployment choices.
For many organizations, the first stage of AI adoption happens in the public cloud. A team needs GPU capacity, provisions an instance, connects storage, and starts experimenting.
Production changes the requirements.
Multiple teams need access. Models and datasets become sensitive. GPU utilization matters. Networking affects training performance. Infrastructure teams need quotas, isolation, identity, monitoring, and a repeatable way to deploy workloads.
At that point, the question is no longer simply where to rent a GPU.
It becomes: how do we operate GPU infrastructure as a cloud?
A private GPU cloud provides dedicated accelerated compute while adding the infrastructure services needed to provision, isolate, govern, and operate that capacity across multiple workloads and teams. OpenStack can provide that cloud foundation, managing GPUs alongside virtual machines, bare metal, networking, and storage, while Kubernetes can orchestrate containerized AI workloads above it.
A private GPU cloud is dedicated GPU infrastructure delivered through a cloud operating model.
Instead of administrators manually assigning individual GPU servers to users, infrastructure becomes a pool of resources that authorized teams can consume through APIs, automation, or self-service workflows.
A rack of GPU servers gives an organization compute capacity.
A private GPU cloud adds the systems required to make that capacity reusable and operational: identity, tenancy, networking, storage, scheduling, quotas, provisioning, monitoring, and lifecycle management.
Depending on the environment, it can run in an organization's own data center, in dedicated hosted infrastructure, or across a combination of locations.
The objective is not simply to own GPUs. It is to make expensive accelerator capacity usable as shared production infrastructure.
Public GPU clouds remain useful, particularly when demand is temporary, experimentation is still early, or teams need access to hardware they do not want to purchase.
Private infrastructure becomes more relevant when AI usage becomes sustained.
Organizations may need tighter control over where datasets and model artifacts reside. Predictable GPU demand can change the economics of continuously rented capacity. Multiple departments may require isolated access to common infrastructure, while production applications may need predictable networking or storage performance.
AI workloads also create requirements beyond the accelerator itself.
Training jobs may need large datasets delivered quickly across several workers. Inference services may require low-latency access to internal applications and data. Research teams may want bare-metal accelerators, while application teams prefer Kubernetes.
The infrastructure therefore has to support several consumption models without turning every AI project into a custom deployment.
That is where the cloud layer becomes important.
OpenStack provides the infrastructure control plane beneath the AI workload.
It can manage compute, networking, storage, identity, and hardware resources through consistent APIs. GPU-enabled resources can then become part of the same infrastructure environment used to provide conventional virtual machines, bare-metal systems, and Kubernetes clusters.
Users should not need to know which physical server contains a particular accelerator before requesting compute. Infrastructure teams can define available resource types and allow the platform to place workloads on appropriate hardware.
OpenStack Nova can expose physical PCI devices such as GPUs through PCI passthrough, while Placement can participate in tracking and scheduling specialized resources.
For some workloads, that may mean an entire physical GPU.
For supported hardware and environments, it may mean a virtualized or partitioned GPU.
For workloads requiring direct access to an entire server, OpenStack Ironic provisions and manages bare-metal machines while integrating with the wider OpenStack environment.
The goal is not to prescribe one way to consume GPUs. It is to give infrastructure teams several controlled, repeatable consumption paths.
VEXXHOST's OpenStack platform can provide the cloud layer for infrastructure spanning virtual machines, bare metal, Kubernetes, storage, networking, and accelerator-enabled workloads.
One of the most common mistakes in AI infrastructure planning is designing around the accelerator first and everything else second.
GPUs are only one part of the system.
Start by defining the workloads the environment actually needs to support.
Large distributed training, fine-tuning, batch processing, development environments, and production inference can have very different requirements.
Some workloads need dedicated physical GPUs. Others can use partitioned or shared resources where supported. Some applications fit comfortably inside virtual machines, while others are designed around containers and Kubernetes.
That makes workload classification more useful than simply selecting the newest GPU available.
The better question is:
What combination of accelerator memory, compute, isolation, network throughput, storage throughput, and availability does this workload require?
Once AI workloads extend across multiple nodes, the network becomes part of the compute path.
Distributed training can generate substantial east-west traffic between GPU workers. Storage systems must move datasets toward compute. Checkpointing can create large bursts of data, while production inference introduces latency-sensitive application paths.
A private AI cloud therefore needs networking designed around its actual workloads rather than simply inheriting whatever architecture already exists in the data center.
Depending on scale and application requirements, that may involve high-bandwidth Ethernet, RDMA, or other high-performance networking technologies.
OpenStack provides software-defined networking and tenant isolation at the infrastructure level, while specialized networking can be integrated where workloads require it.
Fast GPUs waiting for data are expensive idle infrastructure.
Training datasets, checkpoints, model weights, artifacts, and inference data create different storage patterns.
Some workloads require high-throughput shared datasets. Others depend on low-latency block storage. Model repositories and large training archives may be better suited to object storage.
A production architecture should therefore consider storage and GPUs together.
OpenStack can be combined with distributed storage platforms such as Ceph to provide block, object, and file storage while keeping the storage layer independent of proprietary cloud services.
The operating model changes significantly once several teams share a GPU platform.
Infrastructure teams need to answer questions such as:
Private cloud makes those controls part of the infrastructure rather than relying on informal allocation of individual machines.
OpenStack projects, identity policies, quotas, and network isolation can provide boundaries between teams while allowing physical infrastructure to operate as a common pool.
OpenStack and Kubernetes solve different parts of the AI infrastructure problem.
OpenStack manages the infrastructure.
Kubernetes manages containerized workloads running on that infrastructure.
This separation is useful because AI teams increasingly expect Kubernetes-based workflows for training, inference, and machine-learning platforms, while infrastructure teams still need a system for controlling the servers, networks, storage, and accelerators beneath those clusters.
A private GPU cloud can therefore provide Kubernetes clusters on top of OpenStack while retaining a consistent infrastructure control layer underneath them.
A research group may need bare-metal GPU nodes. Another team may want GPU-backed virtual machines. An application platform may consume GPUs through Kubernetes.
The underlying infrastructure can support all three.
There is no single correct GPU allocation model.
Dedicated GPUs provide strong isolation and predictable access to the full accelerator, making them appropriate for demanding training jobs or workloads where performance consistency matters.
Where supported by the GPU, virtualization stack, and vendor technology, virtual GPU or hardware-partitioning capabilities can divide accelerator capacity into smaller resources. This can improve utilization for workloads that do not require an entire device.
The decision should follow workload characteristics rather than an attempt to maximize theoretical utilization.
A development environment with intermittent GPU usage does not necessarily need the same allocation model as a latency-sensitive inference service or multi-node training workload.
The role of the cloud platform is to make those different resource classes repeatable and manageable.
A private AI cloud does not automatically mean the hardware needs to sit inside your own building.
There are two separate decisions:
Where should the infrastructure live?
Who should operate it?
Some organizations require GPU infrastructure inside their own facilities because of data location, existing investments, latency, or organizational policy.
Others want dedicated infrastructure without operating high-density GPU systems themselves. A hosted private GPU cloud can provide dedicated hardware in a provider facility while retaining a private operating model.
Operational responsibility is a separate choice. An organization may operate the platform internally, retain control with specialist support, or use a fully managed model.
VEXXHOST's GPU Infrastructure offering supports dedicated AI infrastructure designed around workload requirements, including hosted and customer-premises deployment models.
Private GPU infrastructure becomes worth evaluating once AI moves from isolated experiments into a recurring organizational capability.
Common indicators include:
It is not automatically the right choice for every AI project.
Teams with sporadic GPU requirements may be better served by public capacity. Early-stage experiments may not justify dedicated infrastructure. Organizations that do not want to operate the platform themselves may prefer hosted or managed infrastructure.
The decision should follow the workload.
One of the most expensive mistakes is purchasing accelerator capacity before designing the system that will keep it useful.
Before committing to hardware, define:
Then define how capacity will be provisioned, maintained, recovered, and returned to the resource pool.
A production environment should be able to take available hardware, expose it through repeatable workflows, monitor it, upgrade it, recover it after failures, and make it available for the next workload.
That is the difference between owning GPU servers and operating GPU infrastructure.
AI infrastructure decisions can persist for years, even while the technology running on top of them changes rapidly.
GPU generations will change. Models and frameworks will change. Kubernetes versions will change.
The infrastructure foundation should make those transitions easier rather than tying the organization to one generation of hardware or one provider-specific operating model.
OpenStack provides a way to treat accelerator resources as part of a broader cloud environment alongside virtual machines, bare metal, networking, storage, identity, and Kubernetes.
VEXXHOST builds GPU infrastructure around workload requirements, with OpenStack and Atmosphere providing the cloud operating layer and deployment options spanning hosted and customer-premises infrastructure.
If your organization is moving from individual GPU servers toward shared production AI infrastructure, start with the workload, data, performance, and operating requirements, not the hardware SKU. Talk to us about open private AI foundation!
Choose from Atmosphere Cloud, Hosted, or On-Premise.
Simplify your cloud operations with our intuitive dashboard.
Run it yourself, tap our expert support, or opt for full remote operations.
Leverage Terraform, Ansible or APIs directly powered by OpenStack & Kubernetes