overfeed.news

Qwen/Qwen3-8B

25d

Age

1.8K

Likes

12.8M

Downloads

Collected

What it costs to run this

8.2B params · 32,768 Context

PrecisionNeedsCheapest that fits
16-bit23GBRTX 3090$0.12/h
8-bit11GBTesla P100$0.03/h
4-bit5.7GBGTX 1660 S$0.02/h

Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. See all GPU prices

New text-generation model. Tags: qwen3, text-generation, conversational, arxiv:2309.00071, arxiv:2505.09388, base_model:Qwen/Qwen3-8B-Base, base_model:finetune:Qwen/Qwen3-8B-Base, license:apache-2.0

Read the full article at huggingface.co

overfeed.news indexes and links. We publish a short excerpt — the full article stays at Hugging Face New Models.

Log in to follow this source
Qwen/Qwen3-8B — overfeed.news