overfeed.news

OpenRouter API Pricing Update: 22 Free Models, Price Cuts for DeepSeek and MoonshotAI, New Paid Entrants

765 words

On 2026-09-12, OpenRouter tracks 443 models with 22 available for free, including new free tiers from Google and Thinking Machines, while paid models saw input price drops for DeepSeek V3 0324 and MoonshotAI Kimi K3.

Free Models

There are 22 free models available on OpenRouter as of 2026-09-12, including Google's Lyria 3 Pro Preview and Lyria 3 Clip Preview, each with 1,048,576 context tokens. Thinking Machines offers two free Inkling variants with the same context length. NVIDIA provides free access to Nemotron 3.5 Lightning and Nemotron 3 Ultra, both with 1,000,000 context tokens. Additional free models come from Dots Studio, inclusionAI, Nex AGI, Poolside, Google (Gemma 4 variants), Cohere, and NVIDIA (Nemotron 3 Super and Nemotron 3 Nano Omni). The Free Models Router from OpenRouter is also available with 200,000 context tokens.

New and Removed Models

Seven new models were added to OpenRouter on 2026-09-12. inclusionAI's Ling 3.0 Flash VL is available at $0.06 per million input tokens and $0.18 per million output tokens with 131,072 context tokens and open weights. Sakana introduced two paid models: Fugu Max at $2.00 input and $6.00 output per million tokens, and Fugu Ultra v2 at $5.00 input and $30.00 output per million tokens, both with 1,000,000 context tokens and closed weights. OpenAI released four new variants under the ~openai provider: GPT Astra Latest ($10.00 input, $50.00 output), GPT Luna Latest ($0.20 input, $1.20 output), GPT Sol Latest ($2.00 input, $10.00 output), and GPT Terra Latest ($2.00 input, $12.00 output), all with 1,050,000 context tokens and closed weights. One model was removed: OpenAI GPT Latest from the ~openai provider.

Price Changes

Several models experienced price adjustments compared to 2026-09-11. MoonshotAI's Kimi K3 saw input drop from $3.00 to $1.6159 and output from $15.00 to $8.1054 per million tokens. DeepSeek V3 0324 reduced input from $0.29 to $0.25 and output from $1.14 to $1.00 per million tokens. Upstage's Solar Pro 4 increased input from $0.03 to $0.09 and output from $0.12 to $0.36 per million tokens. DeepSeek V4 Pro 0813 lowered input from $0.66 to $0.5795 and output from $1.98 to $1.7384 per million tokens. Qwen3 30B A3B Instruct 2507 raised input from $0.0481 to $0.09 and output from $0.193 to $0.30 per million tokens. DeepSeek V4 Flash 0423 cut input from $0.0886 to $0.0676 and output from $0.1772 to $0.1352 per million tokens. Z.ai's GLM 5.2 reduced input from $0.966 to $0.60 and output from $3.036 to $2.00 per million tokens. Conversely, Z.ai's GLM Latest (under ~z-ai) increased input from $0.70 to $0.8727 and output from $2.20 to $3.36 per million tokens. MoonshotAI Kimi Latest (under ~moonshotai) matched Kimi K3's new rates at $1.6159 input and $8.1054 output per million tokens. DeepSeek V4 Pro 0423 made a minor cut, reducing input from $0.87 to $0.8397 and output from $1.74 to $1.6794 per million tokens.

Cheapest Paid Options

The lowest input cost among paid models is IBM's Granite 4.0 Micro at $0.017 per million tokens, with output at $0.112 per million tokens and 131,000 context tokens. Mistral Nemo follows at $0.019 input and $0.03 output per million tokens with 131,072 context tokens. inclusionAI's Ling 3.0 Flash is priced at $0.021 input and $0.063 output per million tokens with 262,144 context tokens. OpenAI's GPT-5 Nano (batch) offers $0.025 input and $0.20 output per million tokens with 400,000 context tokens. Meta's Llama 3.2 1B Instruct charges $0.027 input and $0.201 output per million tokens with 60,000 context tokens. Qwen3.7 Flash, OpenAI's gpt-oss-20b, and Amazon's Nova Micro 1.0 all start at $0.03 input per million tokens, with varying output rates and context lengths.

Largest Context Windows

The models with the largest context windows on OpenRouter are the Auto Router (Beta), Pareto Code Router, and two variants of SpaceXAI's Grok 4.20 (standard and Multi-Agent), all provided by openrouter or x-ai, each offering 2,000,000 context tokens. The standard Auto Router also provides 2,000,000 context tokens. These routers are designed for extended-input use cases, though they are not open-weight models.

  • 22 models are free to use on OpenRouter as of 2026-09-12, with context lengths up to 1,048,576 tokens.
  • Input prices decreased for DeepSeek V3 0324, MoonshotAI Kimi K3, and several DeepSeek V4 variants, while Upstage and some Qwen models saw increases.
  • The cheapest paid model by input cost is IBM's Granite 4.0 Micro at $0.017 per million tokens.
  • Seven new models were added, including four OpenAI GPT variants and inclusionAI's Ling 3.0 Flash VL.
  • The largest available context window is 2,000,000 tokens, offered by multiple routers from OpenRouter and x-ai.
OpenRouter API Pricing Update: 22 Free Models, Price Cuts for…