overfeed.news

pipecat-ai/phonellm-alpha-1

1mo

Age

72

Likes

64

Downloads

Collected

What it costs to run this

31.6B params · 32,768 Context

PrecisionNeedsCheapest that fits
16-bit70GBA800 PCIE$0.47/h
8-bit35GBQ RTX 8000$0.25/h
4-bit17GBTesla P40$0.08/h

Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. See all GPU prices

New text-generation model. Tags: nemotron_h, text-generation, nemotron, mixture-of-experts, voice-agent, phone, tool-use, function-calling

Read the full article at huggingface.co

overfeed.news indexes and links. We publish a short excerpt — the full article stays at Hugging Face New Models.

Log in to follow this source