overfeed.news

LLM API Pricing Update: September 22, 2026

736 words

On September 22, 2026, OpenRouter tracked 443 models with 24 available for free, including new free offerings from Google and Thinking Machines, while paid models saw price adjustments across input and output tokens.

Free Models

As of September 22, 2026, 24 models are available for free on OpenRouter, up from the previous day. These include Google's Lyria 3 Pro Preview and Lyria 3 Clip Preview, both with 1,048,576 context tokens. Thinking Machines offers two free variants: Inkling Small and Inkling, each with 1,048,576 context tokens. NVIDIA provides two free models: Nemotron 3.5 Lightning and Nemotron 3 Ultra, each with 1,000,000 context tokens. Additional free models come from Dots Studio, inclusionAI, Nex AGI, Qwen, Poolside, Google (Gemma 4 variants), Cohere, and NVIDIA (Nemotron 3 Super and Nemotron 3 Nano Omni), with context tokens ranging from 256,000 to 1,048,576.

New and Removed Models

Four new models were added to OpenRouter on September 22, 2026. SpaceXAI's Grok 4.7 is priced at $1.60 per million input tokens and $4.80 per million output tokens with 500,000 context tokens. Xiaomi introduced three models: MiMo-V2.6-Flash at $0.14 input and $0.28 output per million tokens, MiMo-V2.6-Pro at $0.435 input and $0.87 output, and MiMo-V2.6-Pro-UltraSpeed at $4.35 input and $8.70 output, all with 1,048,576 context tokens. Seven models were removed from the platform, including Anthropic's Claude Opus 4, MiniMax M3 (batch), MoonshotAI's Kimi K3 (batch), OpenAI's gpt-oss-120b (batch), and three Qwen batch models, along with Thinking Machines' Inkling (batch).

Price Changes

Several models experienced price adjustments on September 22, 2026. DeepSeek Pro Latest decreased input price from $0.5795 to $0.5586 and output from $1.7384 to $1.6759 per million tokens. DeepSeek Flash Latest lowered input from $0.13 to $0.12 and output from $0.52 to $0.48. MoonshotAI's Kimi K3 increased input from $1.70 to $3.00 and output from $8.50 to $15.00 per million tokens. DeepSeek V4 Flash Latest raised output price from $0.16 to $0.40 while input remained at $0.04. NVIDIA's Nemotron 3 Nano 30B A3B reduced input from $0.06 to $0.05 and output from $0.24 to $0.20. Z.ai's GLM 5.3 Flash increased input from $0.09 to $0.15 and output from $0.30 to $0.50. DeepSeek V4 Pro 0813 cut input from $0.66 to $0.5586 and output from $1.98 to $1.6759. Qwen's Qwen3.8 27B increased input from $0.20 to $0.42 and output from $2.55 to $3.00. DeepSeek V4 Flash 0423 made minor reductions: input from $0.0899 to $0.0886 and output from $0.1797 to $0.1772. Z.ai's GLM Latest increased input from $0.7735 to $0.7839 and output from $2.431 to $2.6532. xAI's Grok Latest lowered input from $2.00 to $1.60 and output from $6.00 to $4.80. MoonshotAI's Kimi Latest decreased input from $1.70 to $1.50 and output from $8.50 to $7.50 per million tokens.

Cheapest Paid Options

The most affordable paid models by input cost are IBM's Granite 4.0 Micro at $0.017 per million input tokens and $0.112 output, with 131,000 context tokens. Mistral Nemo follows at $0.019 input and $0.03 output per million tokens, with 131,072 context tokens. inclusionAI's Ling 3.0 Flash is priced at $0.021 input and $0.063 output per million tokens, offering 262,144 context tokens. OpenAI's GPT-5 Nano (batch) costs $0.025 input and $0.20 output per million tokens with 400,000 context tokens. Meta's Llama 3.2 1B Instruct is at $0.027 input and $0.201 output per million tokens, with 60,000 context tokens. Qwen's Qwen3.7 Flash, OpenAI's gpt-oss-20b, and Inference.net's Schematron V2 Turbo all charge $0.03 per million input tokens, with output prices of $0.13, $0.13, and $0.15 respectively, and context tokens ranging from 128,000 to 1,000,000.

Largest Context Windows

The models with the largest context windows on OpenRouter as of September 22, 2026, are Auto Router (Beta), Pareto Code Router, SpaceXAI's Grok 4.20 Multi-Agent, SpaceXAI's Grok 4.20, and Auto Router, all offering 2,000,000 context tokens. These are provided by OpenRouter and xAI. No other models in the tracked set exceed this context length.

  • Free model availability increased to 24, with new entries from Google and Thinking Machines offering up to 1,048,576 context tokens.
  • Price changes were mixed: DeepSeek and xAI reduced costs, while MoonshotAI and Z.ai increased prices for certain models.
  • The cheapest paid model remains IBM's Granite 4.0 Micro at $0.017 per million input tokens.
  • Xiaomi's MiMo-V2.6-Flash is the lowest-cost new paid model at $0.14 per million input tokens with 1,048,576 context tokens.
  • The largest context windows available are 2,000,000 tokens, held by OpenRouter routers and xAI's Grok 4.20 variants.
LLM API Pricing Update: September 22, 2026 — overfeed.news