overfeed.news

Qwen/Qwen2.5-7B-Instruct

25d

Age

2.1K

Likes

10M

Downloads

Collected

What it costs to run this

7.6B params · 32,768 Context

PrecisionNeedsCheapest that fits
16-bit18GBTesla P40$0.08/h
8-bit9.2GBTesla P100$0.03/h
4-bit4.6GBGTX 1660 S$0.02/h

Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. See all GPU prices

New text-generation model. Tags: qwen2, text-generation, chat, conversational, en, arxiv:2309.00071, arxiv:2407.10671, base_model:Qwen/Qwen2.5-7B

Read the full article at huggingface.co

overfeed.news indexes and links. We publish a short excerpt — the full article stays at Hugging Face New Models.

Log in to follow this source
Qwen/Qwen2.5-7B-Instruct — overfeed.news