overfeed.news

Model fit

Which GPU runs which model

How much video memory a model needs, and the cheapest chip that holds it at today's prices.

10models

32,768Context

Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. Memory figures come from each provider's own listing. We report what they publish and do not correct it.

Llama 3.1 8B

8 B params

16-bit

22GB

Tesla P40$0.08/h

8-bit

11GB

Tesla P100$0.03/h

4-bit

5.4GB

GTX 1660 S$0.02/h

Llama 3.1 70B

70 B params

16-bit

161GB

MI300X$2.39/h

8-bit

81GB

RTX PRO 6000 Max-Q$0.79/h

4-bit

40GB

Q RTX 8000$0.25/h

Llama 3.1 405B

405 B params

16-bit

886GB

MI300X × 5$11.95/h

8-bit

443GB

A800 PCIE × 6$2.80/h

4-bit

221GB

MI350X$5.49/h

Qwen 2.5 7B

7.6 B params

16-bit

18GB

Tesla P40$0.08/h

8-bit

9.1GB

Tesla P100$0.03/h

4-bit

4.6GB

GTX 1660 S$0.02/h

Qwen 2.5 32B

32.5 B params

16-bit

79GB

A800 PCIE$0.47/h

8-bit

39GB

Q RTX 8000$0.25/h

4-bit

20GB

Tesla P40$0.08/h

Qwen 2.5 72B

72.7 B params

16-bit

167GB

MI300X$2.39/h

8-bit

84GB

RTX PRO 6000 Max-Q$0.79/h

4-bit

42GB

Q RTX 8000$0.25/h

Mistral 7B

7.2 B params

16-bit

20GB

Tesla P40$0.08/h

8-bit

10GB

Tesla P100$0.03/h

4-bit

5.0GB

GTX 1660 S$0.02/h

Gemma 2 9B

9.2 B params

16-bit

32GB

RTX 5090$0.21/h

8-bit

16GB

Tesla P100$0.03/h

4-bit

7.9GB

Tesla P100$0.03/h

Gemma 2 27B

27.2 B params

16-bit

71GB

A800 PCIE$0.47/h

8-bit

36GB

Q RTX 8000$0.25/h

4-bit

18GB

Tesla P40$0.08/h

Mixtral 8x7B

46.7 B paramsMoE

16-bit

105GB

MI300X$2.39/h

8-bit

52GB

A800 PCIE$0.47/h

4-bit

26GB

RTX 5090$0.21/h

Mixture of experts: only part of the weights runs per token, but all of them must sit in memory.

Quantizing halves the weights at each step. It also changes output quality, which this page does not measure.

See all GPU prices
Which GPU runs which model — overfeed.news