overfeed.news

inclusionAI/Ling-3.0-flash-Fin

1mo

Age

72

Likes

460

Downloads

Collected

What it costs to run this

127.5B params · 32,768 Context

PrecisionNeedsCheapest that fits
16-bit297GBQ RTX 8000 × 7$1.78/h
8-bit149GBMI300X$2.39/h
4-bit74GBA800 PCIE$0.47/h

Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. See all GPU prices

New text-generation model. Tags: bailing_hybrid, finance, financial-research, agents, tool-use, long-context, mixture-of-experts, text-generation

Read the full article at huggingface.co

overfeed.news indexes and links. We publish a short excerpt — the full article stays at Hugging Face New Models.

Log in to follow this source
inclusionAI/Ling-3.0-flash-Fin — overfeed.news