Editorial flat-vector illustration for a technology blog hero about flagship model pricing: a large rounded model chip in deep blue with a soft cyan glow on a light background, a paper price tag clipped to its corner, three token streams flowing out of the chip into a circular price gauge with small abstract ticks, a small stack of cached blocks beside the gauge, and a slim calendar sheet sliding in from the left. White background with soft blue-and-cyan gradient accents and one subtle warm highlight; minimal flat line icons; clean modern geometric sans-serif typography; no readable text.
Guides & Insights

GPT-5.6 Sol API Pricing: $4 / $20 per 1M Tokens After the Price Cut

Author

Gideon Frost

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

GPT-5.6 Sol costs $4.00 per million input tokens and $20.00 per million output tokens on the GPT-5.6 Sol API as of today — the same pass-through rate the flagship tier carries on the OrcaRouter directory, checked 2026-08-25 at zero markup. On August 24, 2026, OpenAI cut Sol's API price by 20% on input and roughly 33% on output, a temporary promotion that runs through at least November 21, 2026. It is still the most expensive tier in the GPT-5.6 family, and it is the first tier whose list price has moved since the July 30 cut, which skipped the flagship. The levers matter more than the headline: reused input bills at $0.40 per million via prompt caching, the batch API drops to $2.00 / $10.00, and fast mode trades a 2x premium for up to 2.5x speed. This page is the pricing reference for Sol specifically — the exact rate card, what a request actually costs, how the flagship compares to the field, when it is the wrong buy, and how to call it through one OpenAI-compatible endpoint with your own key.

Prices are per million tokens (input / output), as of 2026-08-25. The cut — 20% on input, roughly 33% on output, running through at least November 21, 2026 — is OpenAI's own announcement, reported at face value; the dollar figures are read from OpenAI's published API rates and cross-checked against the OrcaRouter model page for GPT-5.6 Sol. Cache, batch, fast-mode and long-context rates are OpenAI-announced. These are promotional rates, not a permanent reset — prices move, and they can move back; verify before you build.

The exact price, verified today

Here is the full rate card for GPT-5.6 Sol, verified 2026-08-25 at the promotional rates. The same rates apply to requests using the gpt-5.6 alias, and the cuts are rolling out to eligible ChatGPT Work and Codex credit plans; ChatGPT Pro, Plus and Business subscriptions are unchanged.

Price card titled 'GPT-5.6 Sol — API price per 1M tokens', sourced to the OrcaRouter directory checked 2026-08-18 at zero markup. Three highlight cells read Input $5.00, Output $30.00 and Cache read $0.50. Rows below list cache write $6.25, long-context over 272K input $10.00 / $45.00, batch API $2.50 / $15.00, fast mode $10.00 / $60.00, context window 1.05M tokens and max output 128K. Footer: Sol's list price was not cut on July 30 — the cheaper tiers were.

• Input — $4.00 per 1M tokens on a cache miss (down from $5.00, 20% off through November 21).

• Output — $20.00 per 1M tokens (down from $30.00, ~33% off through November 21).

• Prompt cache read — $0.40 per 1M tokens (was $0.50), a 90% discount on any input you resend.

• Prompt cache write — $5.00 per 1M tokens (was $6.25).

• Long-context tier — inputs past roughly 272K tokens bill at $8.00 input / $30.00 output (was $10.00 / $45.00; the same 2x-input, 1.5x-output multiples now apply off the lower base).

• Batch API — $2.00 input / $10.00 output, a flat 50% off (was $2.50 / $15.00), asynchronous, typically ready within about 24 hours.

• Fast mode — $8.00 input / $40.00 output (was $10.00 / $60.00), a 2x premium for up to roughly 2.5x output speed with no change in intelligence.

Context window — ~1.05M tokens, up to 128K output.

OrcaRouter passes the provider rate through with no markup, so the figure on the directory is the figure you pay — and a vendor cut lands on the listing the same day it takes effect. On the live model page the card reads $4.00 / $20.00 per 1M, cache read $0.400, with Vision, Tools, JSON and Reasoning badges. This is the first time Sol's list price has moved since the July 30 reduction, which took GPT-5.6 Luna to $0.20 / $1.20 and GPT-5.6 Terra to $2 / $12 while the flagship held; the August 24 cut is the flagship's turn.

The levers that change the effective price

The base rate is only the starting point. Four modifiers change what Sol actually costs you:

• Prompt caching — reused context bills at $0.40 per million instead of $4 fresh, and cache writes cost $5.00 per million. For any agent or chat that re-sends a large prefix on every turn, caching is the single biggest lever on the bill. OpenAI keys the cache per project and API key, so staying on one key (one channel) is what keeps the hits landing.

• Long-context tier — a request whose input passes roughly 272K tokens moves from the base rate to $8 / $30. A few huge prompts can quietly raise your average price, so batch or truncate when you do not genuinely need the deep context.

• Batch — OpenAI's batch API runs jobs asynchronously with a roughly 24-hour turnaround at exactly half price: $2.00 / $10.00, with cached input at $0.20 and cache writes at $2.50. It stacks with prompt caching and suits evals, bulk classification and offline generation where latency does not matter.

• Fast mode — $8 / $40 for up to ~2.5x output throughput, useful for latency-critical, user-waiting work like live coding assistance. It bills per token at 2x the standard rate, does not count toward purchased Scale Tier TPM packages, and can fall back to standard processing if you surge past the shared rate limit.

What a request actually costs

Per-token prices are abstract, so here are worked examples at today's rates:

• Single chat call, 50K input + 20K output — $0.20 + $0.40 = $0.60 (was $0.85).

• 1M input + 1M output — $4 + $20 = $24 (was $35).

• Agentic turn, 200K cached context + 20K fresh input + 30K output — $0.08 + $0.08 + $0.60 = $0.76. The same turn with a cold cache — 220K fresh input at $0.88 plus $0.60 output — totals $1.48. Caching roughly halves it.

A 10M-token month at a 70% input / 30% output split, with 80% of input served from cache: 1.4M fresh input ($5.60) + 5.6M cached ($2.24) + 3M output ($60) ≈ $68. Uncached, the same month is $88 — and the same month on GPT-5.6 Terra runs about $50, or about $5 on GPT-5.6 Luna. That ~14x spread between Sol and Luna is why tier-matching matters more than any single rate.

Two caveats raise the effective rate. Sol is a reasoning model — internal reasoning tokens bill as output at $20 per million. A hard prompt at a high reasoning-effort setting can spend a large share of its output budget on thinking, so the effective output cost on difficult work runs well above a naive estimate. And OpenAI's tokenizer is more verbose than the previous generation's on some inputs, so the same text can cost more tokens than it would have on GPT-5.5. Tune reasoning effort to the task, cap output where you can, and measure actual usage before projecting a bill.

How Sol pricing compares to the field

Comparison table card titled 'USD per 1M tokens — published list prices, 2026-08-18'. Rows with an input/output column pair per model: GPT-5.6 Sol $5.00/$30.00 highlighted with a 'THIS PAGE' tag; GPT-5.6 Terra $2.00/$12.00; GPT-5.6 Luna $0.20/$1.20; Claude Opus 5 $5.00/$25.00; Claude Fable 5 $10.00/$50.00; Claude Opus 4.8 $5.00/$25.00; Grok 4.6 $2.00/$6.00; DeepSeek V4 Pro $0.44/$0.88 off-peak. Footer: Sol's $30 output is the highest of these flagships except Claude Fable 5, and Grok 4.6 matched Sol's index score at $2/$6.

Read against published list prices on 2026-08-25 (Sol at its promotional rate). Press coverage reads the cut as a response to Anthropic and Chinese-lab pricing pressure — that is third-party interpretation, not an OpenAI statement.

• GPT-5.6 Terra — $2 / $12, the post-cut middle tier. Sol now costs 2x more on input and ~1.7x more on output; Terra is the default when you don't need flagship reasoning.

• GPT-5.6 Luna — $0.20 / $1.20, the post-cut volume tier. A 20x input gap to Sol; Luna is priced for the easy majority of traffic.

Claude Opus 5 — $5 / $25. Sol is now 20% cheaper on both input and output on paper, versus matching Opus 5 on input at the old rates. The comparison already favored Sol before the cut: this blog previously measured Sol at roughly 34.5% fewer tokens than an Anthropic model on the same English/mixed sample, and a published token-economics analysis (TheBlockBeats) estimated Sol's total bill can come in about 21% lower than Claude Opus 5 for equivalent work even when Sol's output rate was higher. Measure your own tokenizer, not the sticker price.

• Claude Fable 5 — $10 / $50. Sol is now 60% cheaper on input and 60% cheaper on output than Anthropic's top tier (was 50% / 40% at the old rates).

• Claude Opus 4.8 — $5 / $25. Sol is now 20% cheaper on input and 20% cheaper on output — the sticker advantage over the previous-generation Anthropic flagship has fully flipped.

Grok 4.6 — $2 / $6. The sharpest number on the board: it matched GPT-5.6 Sol on the Artificial Analysis index (both 61) at 2x cheaper input and ~3.3x cheaper output, down from a 2.5x / 5x gap at the old rates.

DeepSeek V4 Pro — $0.44 / $0.88 off-peak (peak windows double it). A different price universe entirely, for the cheap open-weights lane.

The honest framing: Sol is still a premium flagship, but the cut pulls it out of the top price bracket — $20 output now undercuts Claude Opus 5's $25 and Claude Opus 4.8's $25, leaving Claude Fable 5's $50 as the only closed-flagship output rate above it. It wins on effective cost per unit of hard work (its tokenizer, its reasoning, its long context) rather than on the absolute rate, and the August 24 cut sharpens that argument. If your traffic is short prompts and simple tasks, a $2 Terra, a $0.20 Luna, or a $2 / $6 Grok 4.6 is still the cheaper call — and on the index, Grok 4.6 gives up nothing on paper.

When Sol is the wrong choice

Four cases where this price is not the one you want:

Short, simple prompts. Below roughly 20K tokens with no deep reasoning, the flagship premium is wasted. GPT-5.6 Terra at $2 / $12 or GPT-5.6 Luna at $0.20 / $1.20 handles the volume tier for a fraction of the cost, and a Grok 4.6 or DeepSeek-class model costs less again.

Reasoning-heavy volume. Because thinking bills as output, a hard prompt at high effort can cost well above the nominal output line. If your workload is many difficult calls, measure actual output tokens before committing — the effective rate can surprise you.

Price-per-token as the only goal. Grok 4.6 tied Sol on the index at $2 / $6, and DeepSeek V4 Pro sits at $0.44 / $0.88 off-peak. If a frontier benchmark score isn't the hard requirement, those models deliver far lower per-token rates for the same class of work.

Strict latency budgets. Sol is a large reasoning model, so time-to-first-token runs well above a Flash-tier model's. For real-time chat where sub-second TTFT matters, fast mode at $8 / $40 buys throughput, not lower first-token latency — and a smaller model is still the cheaper fix.

How to call it on OrcaRouter

The OrcaRouter model page for openai/gpt-5.6-sol, showing the $5.00 per million input and $30.00 per million output pricing, the ~1.05M-token context window, 128K max output, and the vision, tools and JSON badges.

GPT-5.6 Sol is hosted on OrcaRouter as model ID openai/gpt-5.6-sol, served through the OpenAI-compatible endpoint at api.orcarouter.ai/v1 at the provider rate with zero markup. If you already have an OpenAI client, migration is a base-URL and model-ID change — nothing else. The same key and endpoint carry 200+ models, so Sol sits next to GPT-5.6 Terra, GPT-5.6 Luna, Claude Opus 5, Claude Fable 5, Grok 4.6 and the open-weight models you'd want to compare it against, and you can route by difficulty instead of re-integrating each one.

Two OrcaRouter specifics worth naming on a pricing page. BYOK: bring your own OpenAI key and OpenAI bills you directly at its own rates — OrcaRouter adds $0 per token and keeps your provider rate limits and credits intact. Guardrails: the PII Shield and content policy run before billing, so a blocked request returns a clean 400 and is never charged, and the Agent Firewall grades tool and MCP calls before they execute.

The honesty note that matters for a pricing page: we route the model, we don't set its price. The $4 / $20 figure is OpenAI's promotional rate passed through — and when OpenAI changed it on August 24, the listing followed the same day, which is exactly why this page is date-stamped. If you call OpenAI directly you get the same rate; OrcaRouter exists for the single-key, zero-markup route if you want one endpoint across the whole model line.

Sources and date

All prices above were verified on 2026-08-25. The GPT-5.6 Sol figures are read from OpenAI's published API pricing, cross-checked against the OrcaRouter model page for GPT-5.6 Sol (checked the same day) and press reporting of the August 24 cut (Indian Express, Economic Times, The Star); batch, fast-mode, cache-write and long-context rates are OpenAI-announced. The Claude, Grok and DeepSeek figures are the current OrcaRouter directory listings, with DeepSeek V4 Pro shown at its off-peak rate. Sol's list price stood unchanged from its July 9 launch through the July 30 cut and moved for the first time on August 24 with this promotion — which is exactly why the date matters here: the $4 / $20 rates are temporary, and prices move; treat any earlier or later "GPT-5.6 Sol pricing" page you find with the date checked. The numbers here are current as of today.

Every price above is the live rate on the GPT-5.6 Sol on OrcaRouter — checked 2026-08-25 at zero markup, ready for your own key.

Compared in this article2

Detected from this article · Benchmarks: Artificial Analysis · updated daily

© 2026 OrcaRouter

For Providers

Run an inference platform? Get your models on OrcaRouter.

providers@orcarouter.ai

Join our community

Discordsupport@orcarouter.aiXGitHubYouTube