overfeed.news

JetBrains/Mellum2.1-12B-A2.5B-Thinking

1h

Age

86

Likes

583

Downloads

Collected

What it costs to run this

12.1B params · 32,768 Context

PrecisionNeedsCheapest that fits
16-bit28GBRTX 5090$0.21/h
8-bit14GBTesla P100$0.03/h
4-bit7.0GBTesla P100$0.03/h

Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. See all GPU prices

New text-generation model. Tags: mellum, text-generation, conversational, en, arxiv:2605.31268, base_model:JetBrains/Mellum2-12B-A2.5B-Base, base_model:finetune:JetBrains/Mellum2-12B-A2.5B-Base, license:apache-2.0

Read the full article at huggingface.co

overfeed.news indexes and links. We publish a short excerpt — the full article stays at Hugging Face New Models.

Log in to follow this source
JetBrains/Mellum2.1-12B-A2.5B-Thinking — overfeed.news