Running a model in production is a different problem from running it in a notebook. The question stops being whether it works and becomes how many requests one GPU can serve and what each one costs.
1. Continuous batching
Serving engines such as vLLM, TGI and SGLang add new requests to a running batch instead of waiting for the whole batch to finish. On chat-style traffic this alone often multiplies throughput several times compared with naive batching.
2. Efficient KV cache management
The key-value cache grows with context length and concurrency, and it is usually what runs out first. Paged attention and prefix caching keep it compact and let requests that share a system prompt reuse the same memory.
3. Quantization
Serving weights in 8-bit or 4-bit formats cuts memory and bandwidth needs, which is exactly what token-by-token generation is limited by. Quality loss is often negligible for 8-bit and acceptable for 4-bit on many tasks; measure it on your own evaluation set.
4. Right-size the GPU
A 7B model on an 80 GB card leaves most of the memory unused. Either serve several models or replicas per GPU, or move small models to cheaper cards and keep large GPUs for large models.
5. Watch cost per token, not cost per hour
Track tokens generated per GPU-hour and divide your hourly rate by it. That single number tells you whether a change helped, and it makes comparing GPU types straightforward.
| Metric | Why it matters |
|---|---|
| Tokens per second per GPU | Throughput you actually sell |
| Time to first token | What users feel as latency |
| GPU memory used by KV cache | Your concurrency limit |
| Cost per million tokens | The number your margin depends on |



