AI token price by provider: current input and output rates
This page compares ai token price across OpenAI, Anthropic, Google, DeepSeek, Grok, and Perplexity, with input and output rates shown side by side, priced per million tokens.
Rates get checked against each provider’s official pricing docs, last verified August 24, 2026.
This is a reference to scan and compare, not a piece to read start to finish.
Token rates shift often (providers cut prices, add new models, adjust tiers), so this page gets kept current rather than treated as a one-time snapshot.
Estimate how much AI tokens are costing your workflows ⇒
Who this AI token price comparison is for
This page is built for anyone who needs to put a real number on API usage before it hits a bill. Four reader types get direct answers here.
Developers estimating API costs
You’re building a feature and need to know what a million calls actually cost before you ship it, not after the invoice lands.
Teams choosing between OpenAI and Anthropic
You’re comparing OpenAI and Anthropic (or Google, DeepSeek, Grok, Perplexity) on rate cards alone, side by side, per million tokens.
Buyers comparing input vs output rates for a workload
Your workload is output-heavy or input-heavy, and the blended “average” price everyone quotes doesn’t tell you what you’ll actually pay.
Buyers weighing cheap large-context models against premium tiers
You’re deciding between a low-cost, large-context model and a premium tier priced for output quality, not just raw token cost.
The rate tables below cover all of that. If you just want your own numbers for your own prompt and response lengths, skip straight to the calculator and get a live estimate in seconds.
AI token price comparison: all providers side by side
Every major provider’s input and output rates sit in the table below, priced per million tokens and checked against provider pricing docs as of August 24, 2026.
Rates change without much warning, so treat this as a snapshot rather than a permanent reference.
Flagship models
| Provider | Model | Input (per 1M) | Output (per 1M) | Context window | Notes |
|---|---|---|---|---|---|
| OpenAI | GPT-5 | $5.00 | $15.00 | 400K | Flagship reasoning model |
| Anthropic | Claude Opus | $15.00 | $75.00 | 200K | Highest-cost tier in the Claude lineup |
| Anthropic | Claude Sonnet | $3.00 | $15.00 | 200K | Mid-tier, most commonly deployed Claude model |
| Gemini 2.5 Pro | $1.25 | $10.00 | 1M | Large context window at flagship tier | |
| DeepSeek | DeepSeek V4 | $0.28 | $1.42 | 128K | Cheapest flagship-class model in this set |
| xAI | Grok 4 | $3.00 | $15.00 | 256K | Priced in line with Claude Sonnet and GPT-5 |
| Perplexity | Sonar Pro | $3.00 | $15.00 | 200K | Search-augmented, adds per-request fee on top |
Budget and lightweight models
| Provider | Model | Input (per 1M) | Output (per 1M) | Context window | Notes |
|---|---|---|---|---|---|
| OpenAI | GPT-5 Mini | $0.25 | $2.00 | 400K | Budget tier for high-volume tasks |
| Anthropic | Claude Haiku | $0.80 | $4.00 | 200K | Fastest, cheapest model in the Claude lineup |
| Gemini 2.5 Flash | $0.30 | $2.50 | 1M | Balances cost and context length | |
| Gemini 2.0 Flash-Lite | $0.075 | $0.30 | 1M | Lowest input rate of any model tracked here | |
| DeepSeek | DeepSeek V4 Lite | $0.14 | $0.56 | 128K | Cut-rate option for simple, high-volume jobs |
Flagship models
The pattern across every provider’s top-tier model is the same: output costs several times more than input, and that gap widens as reasoning depth increases.
Claude Opus sits at the top of the cost curve ($15 input, $75 output per 1M tokens), aimed at complex agentic and coding workloads rather than everyday chat.
GPT-5, Claude Sonnet, and Grok cluster in a similar band, with Google’s Gemini 2.5 Pro standing out for pairing flagship-level output quality with a much larger context window at a lower input rate.
Perplexity’s flagship model matches Sonnet’s per-token rate but adds a separate per-request charge for search retrieval, so its effective cost per query runs higher than the token rate alone suggests.
Budget and lightweight models
Every provider now ships a stripped-down model built for high-volume, low-complexity work, and the price gap between flagship and budget tiers is steep.
Gemini 2.0 Flash-Lite currently posts the lowest input rate of any model tracked here at $0.075 per 1M tokens, useful for classification, tagging, or simple extraction jobs run at scale.
DeepSeek’s lite variant undercuts most competitors on both sides of the ledger, while Claude Haiku and GPT-5 Mini trade a bit more cost for stronger instruction-following.
For workloads that do not need deep reasoning, routing traffic to these lightweight models rather than a flagship tier is usually the single biggest lever for cutting an API bill.
The input-to-output price multiple by model
The output-to-input price multiple, not the headline rate, is what actually decides your cost once a workload has a real shape.
Divide a model’s output price by its input price and you get a number that tells you how punishing long generations are relative to long prompts. Most flagship models sit in a 4x to 6x range. Reasoning and thinking variants often break that pattern entirely, sometimes running 8x to 12x, because their internal reasoning tokens get billed as output even though the user never sees them.
That gap matters more for agentic and reasoning-heavy workloads than the sticker price ever will.
| Model | Input (per 1M) | Output (per 1M) | Output-to-input multiple | Reasoning/thinking variant? |
|---|---|---|---|---|
| Reasoning-tier flagship (example: extended thinking mode) | $3.00 | $36.00 | ~12x | Yes |
| Reasoning-tier flagship (example: o-series style) | $5.00 | $40.00 | ~8x | Yes |
| General flagship model A | $3.00 | $15.00 | ~5x | No |
| General flagship model B | $5.00 | $25.00 | ~5x | No |
| General flagship model C | $2.50 | $10.00 | ~4x | No |
| Lightweight / budget model | $0.15 | $0.60 | ~4x | No |
| Ultra-lightweight model | $0.075 | $0.30 | ~4x | No |
Illustrative multiples based on typical flagship and reasoning-tier pricing patterns. Check current per-provider rates in the comparison table above before budgeting a workload.
A 12x multiple sounds abstract until you put it against a real prompt-to-completion ratio.
If your workload sends short prompts and gets back long, reasoning-heavy answers, that reasoning-tier row dominates your bill regardless of how cheap the input price looked on paper.
Flip the workload (long context stuffed in, short answer out) and a high-multiple model can actually end up cheaper than a low-multiple one with a higher input rate.
This is why the same “cheap model” label breaks down depending on how an agent or pipeline is routed. A chatbot that mostly reads and summarizes behaves nothing like an agent that reasons step by step before replying, even on identical hardware.
Treat the multiple as a screening step before you ever open a per-provider price list: know your expected input-to-output token ratio first, then match it against the multiple, not the sticker price.
Which provider wins your workload shape
The cheapest provider flips depending on whether your workload burns through input tokens or output tokens, because the ranking that wins on flagship price does not automatically win on cost per task.
Two shapes cover most real usage: summarization (long input, short output) and generation or agent loops (short input, long output, repeated many times).
| Workload shape | Example task | Token profile | Cheapest provider/model | Why |
|---|---|---|---|---|
| Summarization | Summarize a 50-page document into a short brief | Long input, short output | Budget/lightweight tier models (e.g. Gemini Flash-class, DeepSeek-class) | Cost is dominated by the input rate since output stays small, so the model with the lowest input price per million tokens wins |
| Generation and agent loops | An agent looping many times, drafting and revising long responses | Short input, long output, repeated calls | Models with the lowest output rate relative to reasoning quality (check the multiple from the table above) | Output tokens get re-fed as input on every loop, so the output rate compounds fast and decides the bill |
Summarization: long input, short output
A summarization task feeds a large document in and asks for a short answer back, which means input pricing decides the cost, not output pricing.
Since the output stays a small fraction of the input, even a model with a steep output rate stays affordable here as long as its input price per million tokens is low.
This is why budget and lightweight models often beat flagship models on this exact shape: their headline cost looks lower, and the workload never touches the expensive part of their rate card.
Generation and agents: short input, long output
An agent loop or long-form generation task starts with a short prompt but produces a long response, and often does that many times in a row.
Here the ranking inverts: the model with the cheapest input rate can end up the most expensive overall if its output rate is high, because output is what gets billed on every pass.
Agent loops make this worse. Each generated token, including tool calls and intermediate steps most users never see, gets appended back into the context and billed again as input on the next turn.
The model that wins this shape is whichever one holds the lowest output rate for the reasoning quality the task needs, not the one with the lowest sticker price on input.
Check the input-to-output multiple from the table above before picking a model for any workload that loops.
Cached input and batch discounts that lower cost
Cached input pricing and batch processing can cut effective token cost well below the headline input and output rates.
For workloads with repeated context (the same system prompt, document, or codebase sent across many requests), caching is often the single biggest lever available, ahead of picking a cheaper model. Batch APIs offer a second, separate discount for anything that doesn’t need a real-time response.
| Provider | Cached input rate | Batch / off-peak discount | Free tier or promo | Notes |
|---|---|---|---|---|
| OpenAI | Reduced rate on repeated prompt prefixes | Batch API, ~50% off standard rate | None ongoing | Caching applies automatically on eligible repeated context |
| Anthropic | Cached reads billed far below fresh input | Batch API, ~50% off (Sonnet 4.6: $1.50 / $7.50) | None ongoing | Cache write costs more than a fresh call, cache read is the savings |
| DeepSeek | Cache hit pricing far below cache miss (V4 Pro: $0.003625 vs $0.435) | Off-peak discount window on top of cache pricing | Low headline rates function as a de facto low-cost tier | Roughly 120x gap between cached and fresh input on repeated context |
| Google Gemini | Context caching available on Pro and Flash tiers | Batch mode discount available | Free API tier with rate limits (Flash-Lite and Flash) | Promo pricing has appeared periodically, confirm current expiry before relying on it |
| Grok | No published cached input rate at time of writing | No published batch discount | None ongoing | Check provider docs directly before assuming a discount applies |
| Perplexity | No published cached input rate at time of writing | No published batch discount | None ongoing | Pricing bundled with search-augmented requests, harder to isolate a pure token rate |
Cached and batch rates matter most for repeated-context workloads: think a chatbot re-sending the same long system prompt on every turn, or a document pipeline reprocessing the same reference file across hundreds of calls.
In those cases, the cache hit rate (not the fresh input rate) becomes your real cost driver, and it can undercut the number on the headline pricing page by an order of magnitude.
Free or low-cost API tiers are genuinely useful for early testing and small-scale experimentation. They still carry rate limits, so treat them as a way to prove out a workflow before committing to a paid tier, not as a production plan.
Promotional pricing, where it exists, tends to carry an expiry date. Confirm the current terms directly with the provider before building a cost model around it.
How to read AI token price rates
AI token price is quoted per million tokens (MTok), with input and output listed as two separate rates.
Output almost always costs more than input, sometimes several times more, because generating tokens takes more compute than reading them.
For the full breakdown of why input and output are metered separately and priced differently, see Input vs Output Tokens in AI: A Simple Guide for Dummies.
One caution worth carrying into the tables above: market-wide “average cost per token” figures get thrown off easily.
A handful of extreme high-price models can pull a provider’s average well above what most workloads actually pay, so treat any single blended number as a rough signal, not a rate you’ll actually see on your invoice.
Estimate your own AI token price with the calculator
Every number in the tables above is a starting point, not your actual bill. Your real cost depends on your specific input length, output length, and model choice, which is exactly what the AI Token Calculator is built to price out.
Price your exact workload, not an average
Stop eyeballing per-million rates. Plug in your real prompt and see the actual cost.
- Input, output, and cached costs across all major models in one view
- Real-time rates for the latest GPT, Claude, Gemini, DeepSeek, and Grok models
- Free, with no signup required
Estimate your AI token costs instantly →
Developers use the free AI Token Calculator and bookmark it to re-check rates whenever providers update pricing.
Bookmarking it takes ten seconds, and it means you are never quoting a client or a budget off memory again.
AI token price frequently asked questions
Which AI API is the cheapest right now?
The cheapest overall option as of August 24, 2026 is a budget model like Gemini Flash-Lite, priced at a fraction of a dollar per million input tokens.
Among flagship models, the cheapest tends to shift by generation, so check the comparison table above for the current leader before you commit to a provider.
Prices move fast enough that naming a permanent winner here would go stale within weeks. Cross-check the rate against the table in this article before you build pricing assumptions into a proposal or budget.
Why are output tokens more expensive than input tokens?
Output tokens cost more because generating text takes more compute than reading it, since every output token runs its own inference step while input gets processed in a single pass.
That gap typically lands at 3 to 5 times the input rate, though it varies by model and provider.
This is why a workload heavy on generation gets expensive fast, even when the headline input price looks cheap.
How does the input to output multiple affect my costs?
The multiple decides your total cost far more than the headline rate once your workload has a real shape.
A model with a cheap input price but a steep output multiple can end up costing more than a pricier-looking model, if your task generates long responses.
Output-heavy work, like drafting, coding, or agent chains, gets dominated by the output rate specifically, not the input rate you saw advertised.
Workloads that mostly read and summarize barely touch the output multiple at all, so the same model can look expensive for one team and cheap for another.
How often do AI token prices change?
AI token prices change often enough that a number checked a few months ago can already be wrong.
Providers cut prices when new model generations ship, and older models sometimes get repriced down or deprecated entirely.
Verify current rates before finalizing a budget or signing off on a vendor. This page and the AI Token Calculator get checked against provider documentation regularly, which is the fastest way to confirm a rate hasn’t shifted since your last look.
Should I pay per token or subscribe to a monthly plan?
Pay per token if your usage is unpredictable or still low volume, since that avoids paying for capacity you won’t use.
A monthly plan tends to make sense once your spend is consistent enough that you can predict it within a reasonable range, often somewhere in the range of a few hundred dollars a month in steady API usage.
Below that threshold, per-token billing keeps you from overcommitting to a plan sized for a workload you don’t have yet.
Do cached tokens really lower my token cost?
Yes, cached tokens lower cost sharply, particularly for workloads that reuse the same context across repeated calls.
A chatbot replaying the same system prompt, or an agent re-reading the same document on every step, sees the biggest drop, since cached input is billed well below the standard input rate.
One-off requests with unique context each time won’t see much benefit, because there’s nothing to cache.