deepseek-ai/DeepSeek-R1
25d
14K
649K
- Collected
What it costs to run this
684.5
| Precision | Needs | Cheapest that fits |
|---|---|---|
| 16-bit | 1528 | MI300X × 8 |
| 8-bit | 764 | RTX PRO 6000 Max-Q × 8 |
| 4-bit | 382 | RTX A6000 × 8 |
Weights plus attention cache, with 15% added for activations and allocator overhead. Everything here is arithmetic — we do not measure speed, so we do not rank chips by it. See all GPU prices
New text-generation model. Tags: deepseek_v3, text-generation, conversational, custom_code, arxiv:2501.12948, license:mit, eval-results, text-generation-inference
overfeed.news indexes and links. We publish a short excerpt — the full article stays at Hugging Face New Models.
More from Hugging Face New Models
Log in to follow this sourceperplexity-ai/pplx-decider-v1.1-27b
New text-classification model. Tags: pytorch, qwen3_5, classification, multimodal, custom-code, decider, decision-model, text-classification
LiquidAI/d1-omni-600M
New image-text-to-text model. Tags: d1_omni, feature-extraction, liquid, lfm2.5, edge, decision, classification, calibration
nerkyor/Qwen3.8-27B-Coder390-EfficientThink-Opus5.5-GPT6Astra-Grok4.7-DSV4Pro-K3-SFT-RLOO-MTP-DFlash2
New image-text-to-text model. Tags: gguf, qwen3.8, efficient-thinking, reasoning, coding, uncensored, sft, simpo
Qwen/Qwen-Image-2.1-Turbo
New text-to-image model. Tags: diffusers, qwen, image-generation, image-editing, text-to-image, base_model:Qwen/Qwen-Image-2.1, base_model:finetune:Qwen/Qwen-Image-2.1, license:other