overfeed.news

incoai/GLM-5.3-Flash-DFlash2

1mo

Age

74

Likes

Collected

What it costs to run this

1.2B params · 32,768 Context

PrecisionNeedsCheapest that fits
16-bit3.2GBGTX 1660 S$0.02/h
8-bit1.6GBGTX 1660 S$0.02/h
4-bit0.8GBGTX 1660 S$0.02/h

Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. See all GPU prices

New text-generation model. Tags: qwen3, dflash, dflash2, speculative-decoding, block-diffusion, draft-model, sglang, text-generation

Read the full article at huggingface.co

overfeed.news indexes and links. We publish a short excerpt — the full article stays at Hugging Face New Models.

Log in to follow this source
incoai/GLM-5.3-Flash-DFlash2 — overfeed.news