STRATA Reserve capacity
Strata/Blog

GPU optimization in Kubernetes: more work from every card

Engineering29 Sep 20267 min read

Kubernetes hands out GPUs one whole card at a time. Most workloads use a fraction of it. Sharing, right-sizing and autoscaling close the gap.

GPU demand behaves differently from CPU demand. A pod that asks for a CPU core and uses half of it wastes a few cents. A pod that asks for a GPU and uses a tenth of it wastes most of an expensive card, every hour it runs.

Why GPUs end up underused

By default the Kubernetes device plugin treats a GPU as an indivisible resource: one pod, one card. That is fine for training, which can saturate a GPU. It is wasteful for notebooks, small models, embeddings and low-traffic inference, which often need only a slice of compute and memory.

Share the card where it's safe

Request what the workload needs

Label nodes by GPU type and memory, and let workloads request the class they need. A 7B inference service does not need an 80 GB card; a fine-tuning job might. Node affinity and separate node pools per GPU type keep expensive cards for the jobs that use them.

Scale on the right signal

CPU-based autoscaling is a poor fit for GPU services. Scale inference on request queue length, tokens per second or GPU utilization exported from DCGM, and scale node pools down when queues are empty. Idle nodes should disappear, not wait.

Measure, then repeat

CheckHealthy sign
GPU utilization per podSteady and high for production, near zero only when scaled to zero
Memory used vs allocatedMost of the card's memory actually in use
Pods pending for GPUsShort waits, no long queues of blocked jobs
Cost per job or per million tokensFalling over time

With Strata you can run these node pools on dedicated bare-metal GPUs and size each pool to its workload, instead of paying for one large, half-idle cluster.

Need GPUs for this?Reserve capacity
More from the blog