GPT-5.6 Sol Price Cuts on Vercel AI Gateway (70% Off Input)
Melvin Vivas · X post · 2026-08-22 · Open on X
Topics: LLMOps, Deployment & Monitoring, LLM Fundamentals · Level: intermediate
Summary
Vercel AI Gateway cut prices again for GPT-5.6 Sol. On top of an existing 50% discount, input is now 20% cheaper and output 33% cheaper, for a total of about 70% off input and 83% off output. Model gateways can be a cheap way to reach frontier models.
Key points
- On top of the 50% discount, input costs 20% less and output 33% less.
- Total savings: about 70% off input and 83% off output.
- The discount covers all tiers, including fast mode.
- Model IDs: openai/gpt-5.6-sol and openai/gpt-5.6-sol-fast.
- Compare model prices across gateways before choosing a provider.
Resources mentioned
- Vercel · tool · x.com · free · recommended by both Bashiri Smith & Melvin Vivas
Deploy web apps and frontends.
Also in: OpenAI DevDay 2026 Recap: Dots, Agents API, Codex Cloud & Marketplace (Melvin Vivas on X · notes), 7 Habits to Become an AI Engineer: Books, Tooling, Research & Shipping (Bashiri Smith on Facebook · notes), Jev model added to the AIBackends API via Vercel AI Gateway (Melvin Vivas on X · notes), Open models now dominate token volume on Vercel AI Gateway (Melvin Vivas on X · notes) and 7 more - GPT 5.6 Sol · tool · openai.com · paid
A GPT-family model available through an API, whose API and credit pricing was cut by over 20% for three months.
Also in: Creator's Top 3 Closed Models: Fable 5, GPT 5.6 Sol, Grok 4.6 (Melvin Vivas on X · notes), Coworker v0.3.1: Telegram Streaming Sync, Built Fast with Codex (Melvin Vivas on X · notes), Coworker: Open-Source Grok Bot Clone Built on Pi and CopilotKit (Melvin Vivas on X · notes), GPT 5.6 Sol (medium) in ChatGPT for planning tasks (Melvin Vivas on X · notes) and 12 more - Vercel changelog: GPT-5.6 Sol is now 50% off at a lower price · article · vercel.com · free
Vercel's changelog post announcing the GPT-5.6 Sol price cuts on AI Gateway.
More in LLMOps, Deployment & Monitoring
- Zero-Shot Prompt Routing with Liquid AI's LFM 2.5-Encoder-350M
- Novita AI Spot GPU Instances: Cheap GPU Compute
- Why Codex Usage Drains Faster: Prompt Cache Hit Rate
- Move Non-Coding Work to Local Models (Hermes + llama.cpp)
- How to Reduce Latency in a Production AI Agent (Interview Answer)
- Serving Ornith-1.5-35B-A3B at 128k Context on an RTX 3090 with llama.cpp