
GPU Infrastructure Upgrades: CUDA, Drivers & Kubernetes
Learn how to safely upgrade production GPU infrastructure across NVIDIA drivers, CUDA, Kubernetes, firmware, containers, and AI workloads.
Lire la noteNotes de terrain / Dernières nouvelles
Des notes d’ingénierie issues de l’exploitation d’infrastructures ouvertes : pannes, décisions de conception et travail upstream qui améliorent l’infrastructure ouverte.
Parcourir toutes les notes
Learn how to safely upgrade production GPU infrastructure across NVIDIA drivers, CUDA, Kubernetes, firmware, containers, and AI workloads.
Lire la note
Your cloud bill is only part of the cost. Learn how operations, licensing, data movement, utilization, and migration affect cloud TCO.
Lire la note
Build an AI infrastructure landing zone for secure, repeatable GPU access across teams with standardized compute, storage, networking, quotas, and operations.
Lire la noteTendances, pratiques exemplaires et analyses techniques approfondies de l’infrastructure infonuagique à code source ouvert.

Velero, Kasten K10, Kanister — the backup tool debate misses the real problem. When your underlying storage changes shape between environments, every backup strategy has to be rebuilt. Here's what a portable infrastructure layer actually fixes.

How Cluster API and OpenStack enable automated cross-region failover with zero downtime. A technical walkthrough using CAPO, Ceph, and Atmosphere.
Hiring takes 6+ months. Your roadmap can't wait. Learn why leading infrastructure teams treat managed services as a permanent layer, not a stopgap.
One open DevOps role triggers overload, burnout, and attrition. See how the cascade runs and how to stop it before the second domino falls.
Learn how to deploy production-grade Kubernetes clusters on bare metal, OpenStack, and public cloud using a single Cluster API-based workflow with Navos. Step-by-step tutorial with manifests.
Data residency is not data sovereignty. The CLOUD Act, CADA, and Canadian procurement policy are forcing a shift toward infrastructure you can control.
Navos runs on unmodified upstream Kubernetes and Cluster API. Learn why that architectural choice matters for portability, multi-cloud, and no-lock-in cluster management.
OpenStack controls infrastructure. Kubernetes orchestrates workloads. For AI, you need both. Learn why the combination delivers what neither can alone.
Hardware cycles are accelerating, regulations are multiplying, and costs are rising. A framework for AI infrastructure decisions that last through 2031.
Learn how mTLS, workload identity, and fine-grained authorization enforce Zero Trust in Kubernetes, a practical guide using Istio, Cilium, and the Navos stack.