overfeed.news

Qwen/Qwen3.8-2.4T-A95B-FP8

2mo

Age

229

Likes

15K

Downloads

Published
Collected

What it costs to run this

2446.2B params · 32,768 Context

PrecisionNeedsCheapest that fits
16-bit5253GBNothing we track holds it
8-bit2627GBNothing we track holds it
4-bit1313GBMI300X × 7$16.73/h

Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. See all GPU prices

New text-generation model. Tags: qwen3_5_moe_text, text-generation, conversational, base_model:Qwen/Qwen3.8-2.4T-A95B, base_model:quantized:Qwen/Qwen3.8-2.4T-A95B, license:other, fp8

Read the full article at huggingface.co

overfeed.news indexes and links. We publish a short excerpt — the full article stays at Hugging Face New Models.

Log in to follow this source
Qwen/Qwen3.8-2.4T-A95B-FP8 — overfeed.news