Two-thirds of AI workloads now run inside Kubernetes clusters. Let that sink in for a second. Two-thirds. Not in fancy custom-built GPU orchestration platforms. Not in some proprietary cloud service with a slick dashboard. In Kubernetes. The same system people were calling „too complex“ for stateless web apps five years ago is now running the most demanding compute workloads on the planet.
I just got back from KubeCon 2026 in Amsterdam, and the keynote was basically a victory lap for this exact thing. And honestly? They earned it. Because I remember when running AI workloads on Kubernetes was pain. Pure, unfiltered pain. You’d spend weeks just getting the GPU drivers to play nice with the container runtime. Now? It’s almost boring how well it works.
We run our AI workloads on Kubernetes too. So let me walk you through what this actually looks like in practice, because the gap between „Kubernetes supports GPUs“ and „I have a model serving traffic in production“ is where all the interesting stuff lives.
The hardware layer: picking your GPU nodes
First things first, you need GPU nodes. We’re running on H100s, which is the sweet spot right now for hosting large language models. But you’ve got options. B200s, B300s, whatever your cloud provider has in stock. The newer generations are faster, obviously, and if you’re serving LLMs, you’ll notice a real difference in tokens per second. We’re talking 2x or 3x improvements depending on the model size and the generation jump.
Here’s the thing people get wrong though. You don’t always need the big guns. If you’re running standard ML workloads, inference on smaller models, classical machine learning stuff, you absolutely do not need an H100. That’s like buying a Ferrari to drive to the grocery store. CPU nodes work fine for a lot of ML tasks. Smaller GPU nodes like the L40S are perfectly capable for medium-sized models. And if you’re on Google Cloud, the TPUs are worth a serious look. They’re genuinely competitive now, not just a Google science project anymore.
The point is: match your hardware to your workload. Don’t blow your entire compute budget on H100s because some blog post told you that you need them. You probably don’t. Unless you’re serving a 70B+ parameter model at scale. Then yeah, get the big nodes.
GPU drivers: the part that used to suck
Once you’ve got GPU nodes in your cluster, you need drivers. This used to be the part where I’d lose a full week of my life and question every career decision I’d ever made. Compatibility matrices, version conflicts, nodes that would randomly stop recognizing their own GPUs after an update. It was brutal.
Now we just use the NVIDIA GPU Operator and it handles everything. Install it, configure it once, forget about it. It manages driver installation, device plugin registration, container toolkit setup, all of it. We’ve been running it for over a year and it has never caused a single problem. Not one. Once you get the initial configuration right, it’s genuinely bulletproof.
The key word there is „once you get the configuration right.“ The first setup took some trial and error with the driver versions and the compatibility flags. But after that? It’s been pure autopilot. This is one of those tools that makes you forget how hard things used to be.
Model serving: where it gets interesting
Okay, so you’ve got your GPU nodes, your drivers are installed, and your cluster is humming. Now you need to actually serve a model. This is where the ecosystem has gotten really exciting.
KServe is probably the most well-known project in this space, and for good reason. It gives you a clean abstraction layer. You define an InferenceService custom resource, point it at your model, and KServe handles the rest. Deployment, scaling, REST API endpoint, health checks, the whole package. It’s genuinely elegant when it works.
But here’s where you have to make a choice. KServe can run in two modes. You can use plain Kubernetes Deployments, or you can go with the Knative plus Istio stack for more advanced features. The Knative route gives you scale-to-zero (which is amazing for cost savings), A/B testing, and canary deployments out of the box. Sounds great, right?
There’s a catch. A big one.
The Knative and Istio stack uses blue-green deployments under the hood. That means when you push an update, it needs to spin up a completely new set of replicas alongside the existing ones before it can cut over. For a web service running on cheap CPU nodes? No big deal. For GPU workloads where each replica occupies an entire H100? You just doubled your GPU requirement for every single deployment. If you have a limited number of GPUs in your cluster, and you almost certainly do, this straight up does not work.
This is why I really prefer the native Kubernetes Deployment mode in KServe. You create a standard Deployment, it creates ReplicaSets, and updates roll out one pod at a time. Classic rolling update. No doubling of resources. No need to keep spare GPUs sitting around just in case you want to deploy a new model version. It just works the way Kubernetes has always worked.
The tradeoff is that you lose scale-to-zero and the fancy traffic splitting. For us, that’s a trade worth making. Your mileage may vary.
The model download problem nobody warns you about
Here’s something that will bite you if you’re not careful: actually getting the model onto your pods.
Maybe you’re pulling from Hugging Face. Maybe from your own S3 bucket or GCS. Maybe from some internal model registry. Wherever it’s coming from, you need to think about rate limits. I’ve seen teams deploy a 10-replica InferenceService and watch every single pod try to download a 40GB model simultaneously from the same source. The download gets throttled, pods timeout during init, they restart, they try to download again, and you end up in this death spiral where nothing ever comes up.
Set your rate limits. Stagger your rollouts. Or better yet, pre-cache your models somewhere close.
And then there’s storage. Where does the model actually live once it’s downloaded? You basically have two options, and they both have problems.
Option one: local disk. Every pod downloads the model to its own ephemeral storage. Simple, but wasteful. If you have ten replicas, you have ten copies of the same model burning through disk space. Plus if a pod gets rescheduled, it has to download the whole thing again.
Option two: network file storage (NFS). One copy of the model, shared across all pods. Much more efficient. But now you’re dependent on network throughput for every single inference request that needs to load model weights, and NFS performance under heavy concurrent access can be… let’s call it „variable.“
We went with NFS and it’s been fine for our use case. But you need to size your NFS throughput correctly and test it under realistic load before you go to production. Don’t learn this lesson the hard way.
So why Kubernetes? Why all this complexity?
I can hear you asking it. „This sounds like a lot of work. Why not just use VMs? Why not just use a managed service?“
And look, you’re not wrong that there’s complexity here. You need to configure pod disruption budgets so your model doesn’t get evicted during a node drain. You need topology spread constraints so all your replicas don’t end up on the same node. You need security contexts, network policies, maybe Falco for runtime security monitoring. You might want dedicated node pools with taints to keep your GPU workloads isolated from the rest of your cluster.
But here’s the thing: you have all the same problems with VMs. You still need load balancing. You still need rolling updates. You still need security. You still need monitoring. You still need to handle node failures gracefully. The difference is that with VMs, you’re building all of that yourself. With Kubernetes, it’s already there. Battle-tested. Used by literally every major tech company on the planet.
Once you have everything set up correctly, deploying a new model is one YAML file. That’s it. You apply a custom resource and Kubernetes pulls the model from your object storage, spins up the pods, configures the networking, and starts serving traffic. You don’t think about how to configure the load balancer. You don’t think about how to do a rolling update. You don’t think about what happens when a node goes down. It’s all solved. The abstractions are good. They’re proven. And they work.
For me, Kubernetes is still the way to go for AI workloads. Not because it’s simple. It’s not. But because the complexity is front-loaded. You pay the setup cost once, and then you get to move fast forever. And in a world where every company is trying to ship AI features yesterday, moving fast is the whole game.
Peace, nerds.