overfeed.news

OpenRouter API Pricing Update: 24 Free Models, Price Increases for DeepSeek V3 and Others

596 words

On 2026-09-18, OpenRouter tracks 445 models with 24 available for free, including new free offerings from Qwen and Thinking Machines.

Free Models

There are 24 free models available on OpenRouter as of 2026-09-18, including Google's Lyria 3 Pro Preview and Lyria 3 Clip Preview, both with 1,048,576 context tokens.

Thinking Machines offers two free models: Inkling Small and Inkling, each with 1,048,576 context tokens. NVIDIA provides two free models: Nemotron 3.5 Lightning and Nemotron 3 Ultra, each with 1,000,000 context tokens.

Additional free models come from Dots Studio, inclusionAI (three variants of Ling 3.0 Flash), Nex AGI (two variants of Nex-N2.5), Qwen (Qwen3.8 27B), Poolside (two Laguna variants), Google (two Gemma 4 variants), NVIDIA (Nemotron 3 Super), Cohere (North Mini Code), and NVIDIA (Nemotron 3 Nano Omni).

New and Removed Models

Two new models were added: Qwen: Qwen3.8 27B (free) from provider qwen, and Pareto from provider unbiased, priced at $2.50 per million input tokens and $7.50 per million output tokens, both with 262,144 context tokens.

One model was removed: Union Alpha from provider stealth. No models became free or stopped being free compared to the previous day.

Price Changes

Eight models had price changes. DeepSeek: DeepSeek V3 (provider deepseek) increased input price from $0.2574 to $0.32 per million tokens and decreased output price from $1.0287 to $0.89 per million tokens.

NVIDIA: Nemotron 3 Nano 30B A3B (provider nvidia) increased input from $0.05 to $0.06 and output from $0.20 to $0.24 per million tokens. Z.ai: GLM 5.3 Flash (provider z-ai) increased input from $0.07 to $0.09 and output from $0.2333 to $0.30 per million tokens.

Z.ai: GLM Flash Latest (provider ~z-ai) increased input from $0.07 to $0.075 and output from $0.2333 to $0.25 per million tokens. NVIDIA: Nemotron 3 Ultra (provider nvidia) increased input from $0.60 to $0.625 and output from $2.40 to $3.125 per million tokens.

Meta: Muse Glimmer 30B (provider meta) decreased input from $0.35 to $0.30 and output from $1.50 to $1.10 per million tokens. Z.ai: GLM Latest (provider ~z-ai) increased input from $0.833 to $0.8775 and output from $2.618 to $2.97 per million tokens.

MoonshotAI: Kimi Latest (provider ~moonshotai) increased input from $2.00 to $2.10 and decreased output from $11.20 to $10.95 per million tokens.

Cheapest Paid Options

The cheapest paid model by input price is IBM: Granite 4.0 Micro (provider ibm-granite) at $0.017 per million input tokens and $0.112 per million output tokens, with 131,000 context tokens and open weights.

Second is Mistral: Mistral Nemo (provider mistralai) at $0.019 per million input and $0.03 per million output, with 131,072 context tokens and open weights.

Third is inclusionAI: Ling 3.0 Flash (provider inclusionai) at $0.021 per million input and $0.063 per million output, with 262,144 context tokens and open weights.

Largest Context Windows

The largest context windows available are 2,000,000 tokens, held by five models: Auto Router (Beta), Pareto Code Router, and two variants of SpaceXAI: Grok 4.20 Multi-Agent and Grok 4.20 (all from provider x-ai), and Auto Router (provider openrouter). All are provided by openrouter or x-ai.

  • 24 models are free to use on OpenRouter as of 2026-09-18, including newly free Qwen3.8 27B.
  • DeepSeek V3 saw the largest input price increase among changed models, rising to $0.32 per million tokens.
  • IBM Granite 4.0 Micro remains the cheapest paid model at $0.017 per million input tokens.
  • No models changed their free status compared to yesterday; all free models were already free on 2026-09-17.
  • The maximum context window available is 2,000,000 tokens, offered by multiple routing and proprietary models.
OpenRouter API Pricing Update: 24 Free Models, Price Increases for…