GPU demand behaves differently from CPU demand. A pod that asks for a CPU core and uses half of it wastes a few cents. A pod that asks for a GPU and uses a tenth of it wastes most of an expensive card, every hour it runs.
Why GPUs end up underused
By default the Kubernetes device plugin treats a GPU as an indivisible resource: one pod, one card. That is fine for training, which can saturate a GPU. It is wasteful for notebooks, small models, embeddings and low-traffic inference, which often need only a slice of compute and memory.
Share the card where it's safe
- Multi-Instance GPU (MIG): on A100, H100 and newer data-center GPUs, one card can be split into isolated instances with their own memory and compute. Good for many small, steady workloads that need predictable performance.
- Time-slicing: several pods take turns on one GPU. There is no memory isolation, so it suits development and bursty jobs, not production with strict latency targets.
- Multi-model serving: run several small models in one inference server process instead of one pod per model.
Request what the workload needs
Label nodes by GPU type and memory, and let workloads request the class they need. A 7B inference service does not need an 80 GB card; a fine-tuning job might. Node affinity and separate node pools per GPU type keep expensive cards for the jobs that use them.
Scale on the right signal
CPU-based autoscaling is a poor fit for GPU services. Scale inference on request queue length, tokens per second or GPU utilization exported from DCGM, and scale node pools down when queues are empty. Idle nodes should disappear, not wait.
Measure, then repeat
| Check | Healthy sign |
|---|---|
| GPU utilization per pod | Steady and high for production, near zero only when scaled to zero |
| Memory used vs allocated | Most of the card's memory actually in use |
| Pods pending for GPUs | Short waits, no long queues of blocked jobs |
| Cost per job or per million tokens | Falling over time |
With Strata you can run these node pools on dedicated bare-metal GPUs and size each pool to its workload, instead of paying for one large, half-idle cluster.



