STRATA Reserve capacity
Strata/Blog

Running GPU workloads on Kubernetes: what it does and what it doesn't

Engineering23 Sep 20266 min read

Kubernetes places GPU jobs. It does not tell you whether they needed that GPU. Knowing the difference saves money and debugging time.

Kubernetes was not designed with GPUs in mind, but it has become the default way to run them. It is very good at some parts of the job and silent about others.

What Kubernetes does well

What it doesn't do

Fill the gaps

GapTypical tool
GPU metricsNVIDIA DCGM exporter with Prometheus and Grafana
GPU sharingMIG or time-slicing in the GPU Operator
Autoscaling on GPU loadKEDA or custom metrics on queue length and utilization
Batch and multi-node jobsKueue, Volcano or a training operator

Where Strata fits

Strata provides the layer Kubernetes can't: the GPU nodes themselves. Join dedicated bare-metal H100, H200 or B200 servers to your cluster, and when a card fails we replace the node, so your scheduler keeps working with healthy capacity.

Need GPUs for this?Reserve capacity
More from the blog