
Shared vs. Dedicated GPUs for Enterprise AI
Compare dedicated GPUs, MIG, vGPU, and time-slicing for enterprise AI. Learn how isolation, performance, utilization, and workload type affect GPU allocation.
Read field noteField notes / Latest
Engineering notes from operating open infrastructure: the failures, design decisions, and upstream work that make open infrastructure better.
Browse all field notes
Compare dedicated GPUs, MIG, vGPU, and time-slicing for enterprise AI. Learn how isolation, performance, utilization, and workload type affect GPU allocation.
Read field note
Why infrastructure flexibility matters and how open cloud technologies can help organizations adapt as workloads and business needs change.
Read field note
Learn how AI workloads change network design across GPU clusters, storage, east-west traffic, RDMA/RoCE, Kubernetes placement, and production inference.
Read field noteAPI vs. self-hosting" is the wrong framing—who operates the model and where it runs are two separate decisions. A guide to the full grid of options, and how to compare their real costs.
Teams evaluating this usually frame it as two options: keep paying a per-token API, or stand up your own serving stack. That framing hides the choice that matters, because it merges two independent decisions. Who operates the serving path is one question. Where the infrastructure sits is a separate one, and the combinations that result include several options most evaluations never price.
The cell most often missing from evaluations is customer-operated serving on rented capacity. Self-operating a model does not require buying accelerators, and treating those as the same decision inflates the apparent cost of control. It also obscures the reverse option, where a provider operates a managed endpoint on hardware you own, which resolves several placement constraints without adding a serving stack to your team's responsibilities.

What forces a change
Unit economics at volume. Per-token pricing is well suited to volume that is small or uncertain. What it buys is broader than elasticity: model access, serving operations, availability, model updates, and capacity that someone else provisions. The question is whether your workload still needs all of those bundled once volume becomes large and predictable.
Data handling. Exposure depends on which cell you are in, and the differences are contractual as much as technical. A public shared endpoint sends prompts to provider-controlled infrastructure under the provider's terms, where retention windows, training-use clauses, and subprocessor lists govern what happens next. A managed service running inside your own cloud tenancy sits in a different contractual position. Private networking and dedicated deployments narrow exposure further. An endpoint on your own premises keeps prompt content in your facility.
Self-hosting does not resolve these requirements on its own. Whether any configuration satisfies a privacy, residency, or compliance obligation depends on the architecture, who holds operator access, where backups and logs are written and retained, whether support access is possible and how it is audited, which subprocessors are involved, and what the model licence permits. Open weights carry licence terms that vary considerably, and some restrict commercial use or redistribution in ways that matter for a production service.
Jurisdiction. Processing location is answerable in several cells, not just the on-premises one, which is why establishing the requirement precisely tends to widen the options rather than narrow them to one.
Model control. Fine-tuned or custom weights, pinning against a provider's deprecation schedule, and evaluating a replacement without rewriting the application all argue for control over the serving path. That is the serving axis, and it can be satisfied on rented capacity.
A per-token rate and a GPU-hour rate are not comparable figures, and dividing one by the other produces a crossover point that will not survive contact with production. A defensible comparison holds the following constant across every option under consideration.
Express the result as cost per successful business task rather than cost per token or per GPU-hour. A task often takes several calls, retries change the arithmetic, and a cheaper model that fails a quality threshold more frequently can cost more per completed piece of work than an expensive one that succeeds first time.
Run a representative workload against each shortlisted option and record the following. NVIDIA's benchmarking documentation makes the necessary warning explicit: tool implementations vary, so results should be compared only when the definitions align. Use one tool, one workload profile, and one set of definitions across every option.

Report percentiles rather than averages throughout. Mean latency conceals the behavior that produces complaints.
Establish the binding constraint first. Cost, data handling, jurisdiction, and model control have different solutions, and only some of them require changing who operates the serving path. Teams that need residency and conclude they must self-operate have skipped the placement axis.
Then ask who carries the pager. A serving tier in a user-facing request path needs the same rotation as anything else that can take a product down. Where that rotation does not exist, provider-managed serving is the answer regardless of what the cost model suggests, and the placement axis is still available for residency and data-handling requirements.
AI Inference covers the provider-managed row of that table. VEXXHOST operates a managed endpoint for an open, commercial, or custom model, deployed in VEXXHOST infrastructure or on your premises. Residency can be shaped around jurisdictional requirements, Zero Data Retention is available for eligible deployment configurations, and pricing is customized to the selected model, expected token usage, and deployment location. The application interface stays stable while model and deployment choices change beneath it, which matters when model control was the reason for moving.
The engagement starts by defining the task, quality threshold, latency, volume, data sensitivity, and application context, then shortlisting models and deployment paths against quality, cost, licensing, and governance. That sequence is the pilot described above, run before committing to a platform.
Where the chosen configuration calls for dedicated hosted capacity or deployment onto hardware in your own facility, GPU Infrastructure provides that layer, sized against the workload's memory, throughput, latency, and availability requirements, with storage and network fabric designed as part of the system.
Start by writing down the binding constraint and the quality threshold. Those two determine which row and column apply, and neither requires a vendor conversation to establish.
Choose from Atmosphere Cloud, Hosted, or On-Premise.
Simplify your cloud operations with our intuitive dashboard.
Run it yourself, tap our expert support, or opt for full remote operations.
Leverage Terraform, Ansible or APIs directly powered by OpenStack & Kubernetes