Kubernetes was not designed with GPUs in mind, but it has become the default way to run them. It is very good at some parts of the job and silent about others.
What Kubernetes does well
- Placement: with the NVIDIA device plugin, pods that request GPUs land on nodes that have them.
- Isolation and scheduling: namespaces, quotas and priorities decide who gets GPUs first.
- Recovery: failed pods are restarted and rescheduled on healthy nodes.
- Repeatability: the same manifests run in staging and production.
What it doesn't do
- Know whether a GPU is used well. A pod at 5% utilization looks the same to the scheduler as one at 95%.
- Understand GPU topology by default. Multi-GPU jobs care which GPUs share NVLink and which node they sit on; that needs extra configuration.
- Scale on GPU signals. Out of the box, autoscaling looks at CPU and memory, not tokens per second or queue length.
- Manage the hardware underneath. Drivers, firmware, failing cards and node replacement are outside its view.
Fill the gaps
| Gap | Typical tool |
|---|---|
| GPU metrics | NVIDIA DCGM exporter with Prometheus and Grafana |
| GPU sharing | MIG or time-slicing in the GPU Operator |
| Autoscaling on GPU load | KEDA or custom metrics on queue length and utilization |
| Batch and multi-node jobs | Kueue, Volcano or a training operator |
Where Strata fits
Strata provides the layer Kubernetes can't: the GPU nodes themselves. Join dedicated bare-metal H100, H200 or B200 servers to your cluster, and when a card fails we replace the node, so your scheduler keeps working with healthy capacity.



