OpenRouter API Pricing Update: September 16, 2026
On 2026-09-16, OpenRouter tracks 443 models with 23 free models, including two new Google Lyria 3 Previews and Thinking Machines Inkling variants, while Qwen Qwen3.5-35B-A3B input price dropped to $0.1625 per million tokens.
Free Models
As of 2026-09-16, there are 23 free models on OpenRouter, including Google: Lyria 3 Pro Preview and Google: Lyria 3 Clip Preview, each with 1,048,576 context tokens. Thinking Machines offers two free models: Inkling Small (free) and Inkling (free), both with 1,048,576 context tokens. NVIDIA provides two free models: Nemotron 3.5 Lightning (free) and Nemotron 3 Ultra (free), each with 1,000,000 context tokens.
Additional free models include Dots Studio: Dots3-Note Preview (free) with 512,000 context tokens, and several inclusionAI Ling 3.0 Flash variants (Sante, Fin, VL) each with 262,144 context tokens. Nex AGI offers Nex-N2.5-Mini and Nex-N2.5-Pro (free), both with 262,144 context tokens. Poolside provides Laguna S 2.1 and Laguna XS 2.1 (free), each with 262,144 context tokens. Google offers Gemma 4 26B A4B and Gemma 4 31B (free), each with 262,144 context tokens. NVIDIA's Nemotron 3 Super (free) has 262,144 context tokens. Cohere's North Mini Code (free) has 256,000 context tokens. NVIDIA's Nemotron 3 Nano Omni (free) has 256,000 context tokens. The Free Models Router from openrouter has 200,000 context tokens.
Model Additions and Removals
One new model was added: Z.ai: GLM 5.2 (free), with input and output prices at $0.0 per million tokens, 32,768 context tokens, and open weights. Three models were removed: Google: Gemma 4 31B (batch), OpenAI: gpt-oss-20b (batch), and Thinking Machines: Inkling Small (batch). No models became free or stopped being free on this date.
Price Changes
Five models had price changes. Qwen: Qwen3.5-35B-A3B input price decreased from $0.3125 to $0.1625 per million tokens, while output price increased from $1.25 to $1.30 per million tokens. Z.ai: GLM 5.3 Flash input price dropped from $0.15 to $0.075 per million tokens, and output price fell from $0.50 to $0.25 per million tokens. Z.ai: GLM 5.2 input price rose from $0.6832 to $1.40 per million tokens, and output price increased from $2.1472 to $4.40 per million tokens. Z.ai: GLM Latest input price decreased from $0.90 to $0.85 per million tokens, and output price fell from $3.00 to $2.8985 per million tokens. DeepSeek: DeepSeek V4 Flash 0731 input price decreased from $0.06 to $0.055 per million tokens, and output price dropped from $0.12 to $0.11 per million tokens.
Cheapest Paid Options
The cheapest paid model by input price is IBM: Granite 4.0 Micro at $0.017 per million tokens for input and $0.112 for output, with 131,000 context tokens and open weights. Second is Mistral: Mistral Nemo at $0.019 input and $0.03 output per million tokens, 131,072 context tokens, open weights. Third is inclusionAI: Ling 3.0 Flash at $0.021 input and $0.063 output per million tokens, 262,144 context tokens, open weights. Fourth is OpenAI: GPT-5 Nano (batch) at $0.025 input and $0.20 output per million tokens, 400,000 context tokens, not open weights. Fifth is Meta: Llama 3.2 1B Instruct at $0.027 input and $0.201 output per million tokens, 60,000 context tokens, open weights.
Largest Context Windows
The largest context windows are held by five models, each with 2,000,000 tokens: Pareto Code Router (openrouter), Auto Router (Beta) (openrouter), SpaceXAI: Grok 4.20 Multi-Agent (x-ai), SpaceXAI: Grok 4.20 (x-ai), and Auto Router (openrouter).
Key takeaways
- Free model count remains at 23, with new additions from Z.ai and no changes in free status for existing models.
- Qwen Qwen3.5-35B-A3B saw the largest input price drop among changed models, halving to $0.1625 per million tokens.
- IBM Granite 4.0 Micro is the cheapest paid model at $0.017 per million input tokens.
- Five models share the maximum context window of 2,000,000 tokens, all from OpenRouter or xAI providers.
- Z.ai: GLM 5.2 experienced the largest price increase, with output rising to $4.40 per million tokens.