overfeed.news

orcarouter/OrcaSAQ-2-27B

14d

Age

112

Likes

359

Downloads

Collected

What it costs to run this

6.8B params · 32,768 Context

PrecisionNeedsCheapest that fits
16-bit24GBRTX 3090$0.12/h
8-bit12GBTesla P100$0.03/h
4-bit5.9GBGTX 1660 S$0.02/h

Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. See all GPU prices

New text-generation model. Tags: vllm, qwen3_5, qwen, qwen3.8, orcasaq2, quantization, mixed-precision, 3-bit

Read the full article at huggingface.co

overfeed.news indexes and links. We publish a short excerpt — the full article stays at Hugging Face New Models.

Log in to follow this source
orcarouter/OrcaSAQ-2-27B — overfeed.news