Notes on GPUs, pricing and running AI in production.

How to choose a GPU for your AI workload
H100, H200, B200 or an RTX 5090? The right card depends less on benchmarks and more on three numbers: model size, batch size and how long the job runs.
Read · 7 min →
On-demand or reserved GPUs: when to commit
On-demand capacity is flexible and expensive per hour. Reserved capacity is cheaper and inflexible. The break-even depends on one number most teams never calculate.
Read · 6 min →
Measure before you scale: find your real bottleneck
Adding GPUs to an I/O-bound job buys you a more expensive version of the same problem. Four measurements tell you which axis to scale.
Read · 6 min →
Start without a GPU: what to validate first
The most expensive GPU hours are the ones spent finding bugs a laptop could have caught. Here is what to check before you rent anything.
Read · 5 min →
Getting more inference out of every GPU
Most production GPUs serving language models run far below their capacity. Five techniques close the gap without changing the model.
Read · 7 min →
GPU rental prices in 2026: what moved and why
H100 rates fell to $1.70 an hour in late 2025 and climbed back above $3 by September 2026. What that swing means for anyone buying compute.
Read · 6 min →
GPU optimization in Kubernetes: more work from every card
Kubernetes hands out GPUs one whole card at a time. Most workloads use a fraction of it. Sharing, right-sizing and autoscaling close the gap.
Read · 7 min →
Idle GPU capacity: the cost nobody sees
The gap between the GPU hours you pay for and the GPU work you get is often bigger than any price difference between providers.
Read · 6 min →
Put a human where GPU spend changes
GPU cost overruns rarely come from one bad decision. They come from spend that grows by default. A few approval points fix that without slowing teams down.
Read · 5 min →
Running GPU workloads on Kubernetes: what it does and what it doesn't
Kubernetes places GPU jobs. It does not tell you whether they needed that GPU. Knowing the difference saves money and debugging time.
Read · 6 min →