Most teams pick a GPU the way people pick a laptop: they look at the most expensive option, flinch, and step down one tier. A better approach starts from the workload. Three questions settle most decisions.
1. Does the model fit in memory?
Memory is the hard limit. A model in 16-bit precision needs roughly two bytes per parameter for the weights alone, so a 70B-parameter model needs about 140 GB before you count activations, optimizer state or the KV cache for inference. Training needs several times more than inference because gradients and optimizer states live in memory too.
- Up to ~13B parameters, inference or LoRA: a 24–32 GB card such as the RTX 4090 or RTX 5090 is usually enough, especially with 8-bit or 4-bit quantization.
- 30–70B parameters, inference: an 80 GB H100 or A100, or a 141 GB H200 if you want long context and large batches on a single card.
- Full fine-tuning or pretraining: multi-GPU nodes with NVLink, typically 8×H100, 8×H200 or B200.
2. Is the job bound by compute, memory bandwidth or the network?
Large-batch training is mostly compute-bound, so raw tensor throughput matters and newer generations pay off. Token-by-token LLM inference is often memory-bandwidth-bound: the GPU spends its time reading weights, not doing math. That is why the H200, with the same compute as the H100 but more and faster memory, can serve noticeably more tokens per second on large models. Multi-node training adds a third constraint: the interconnect between GPUs and between nodes.
3. How long will it run?
A two-hour experiment and a six-month production service are different purchases. For short jobs, the cheapest card that fits is usually the right one, even if it is slower, because you only pay for the hours you use. For long-running services, cost per useful unit of work matters more than the hourly rate: a faster card that serves twice the tokens at 1.5× the price is cheaper per token.
A quick decision table
| Workload | Start with | Move up when |
|---|---|---|
| Image and video generation | RTX 5090 | Video models run out of 32 GB |
| LoRA fine-tuning up to 13B | RTX 5090 or L40S | Batch size or context gets too small |
| LLM inference, 30–70B | H100 | Long context or high concurrency: H200 |
| Full fine-tuning, pretraining | 8×H100 | Time to result matters more than price: H200 or B200 |
The mistake to avoid
Do not size for the peak you might hit next year. Rent what fits the workload today, measure utilization, and move up when the numbers tell you to. With rented capacity, changing GPUs is a configuration change, not a procurement project.



