Model fit
Which GPU runs which model
How much video memory a model needs, and the cheapest chip that holds it at today's prices.
10
32,768
Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. Memory figures come from each provider's own listing. We report what they publish and do not correct it.
Llama 3.1 8B
22
Tesla P40$0.08/h
11
Tesla P100$0.03/h
5.4
GTX 1660 S$0.02/h
Llama 3.1 70B
161
MI300X$2.39/h
81
RTX PRO 6000 Max-Q$0.79/h
40
Q RTX 8000$0.25/h
Llama 3.1 405B
886
MI300X × 5$11.95/h
443
A800 PCIE × 6$2.80/h
221
MI350X$5.49/h
Qwen 2.5 7B
18
Tesla P40$0.08/h
9.1
Tesla P100$0.03/h
4.6
GTX 1660 S$0.02/h
Qwen 2.5 32B
79
A800 PCIE$0.47/h
39
Q RTX 8000$0.25/h
20
Tesla P40$0.08/h
Qwen 2.5 72B
167
MI300X$2.39/h
84
RTX PRO 6000 Max-Q$0.79/h
42
Q RTX 8000$0.25/h
Mistral 7B
20
Tesla P40$0.08/h
10
Tesla P100$0.03/h
5.0
GTX 1660 S$0.02/h
Gemma 2 9B
32
RTX 5090$0.21/h
16
Tesla P100$0.03/h
7.9
Tesla P100$0.03/h
Gemma 2 27B
71
A800 PCIE$0.47/h
36
Q RTX 8000$0.25/h
18
Tesla P40$0.08/h
Mixtral 8x7B
MoE105
MI300X$2.39/h
52
A800 PCIE$0.47/h
26
RTX 5090$0.21/h
Mixture of experts: only part of the weights runs per token, but all of them must sit in memory.
Quantizing halves the weights at each step. It also changes output quality, which this page does not measure.
See all GPU prices