STRATA Reserve capacity
Strata/Blog

Getting more inference out of every GPU

Inference1 Oct 20267 min read

Most production GPUs serving language models run far below their capacity. Five techniques close the gap without changing the model.

Running a model in production is a different problem from running it in a notebook. The question stops being whether it works and becomes how many requests one GPU can serve and what each one costs.

1. Continuous batching

Serving engines such as vLLM, TGI and SGLang add new requests to a running batch instead of waiting for the whole batch to finish. On chat-style traffic this alone often multiplies throughput several times compared with naive batching.

2. Efficient KV cache management

The key-value cache grows with context length and concurrency, and it is usually what runs out first. Paged attention and prefix caching keep it compact and let requests that share a system prompt reuse the same memory.

3. Quantization

Serving weights in 8-bit or 4-bit formats cuts memory and bandwidth needs, which is exactly what token-by-token generation is limited by. Quality loss is often negligible for 8-bit and acceptable for 4-bit on many tasks; measure it on your own evaluation set.

4. Right-size the GPU

A 7B model on an 80 GB card leaves most of the memory unused. Either serve several models or replicas per GPU, or move small models to cheaper cards and keep large GPUs for large models.

5. Watch cost per token, not cost per hour

Track tokens generated per GPU-hour and divide your hourly rate by it. That single number tells you whether a change helped, and it makes comparing GPU types straightforward.

MetricWhy it matters
Tokens per second per GPUThroughput you actually sell
Time to first tokenWhat users feel as latency
GPU memory used by KV cacheYour concurrency limit
Cost per million tokensThe number your margin depends on
Need GPUs for this?Reserve capacity
More from the blog