LLM Cost Calculator.
Enter the tokens in a typical request and how many you send a day. You'll see what each current Claude, GPT and Gemini model would cost, cheapest first.
Your workload
Cost by model
Standard API prices, checked October 8, 2026. A month is 30 days.
| Model | Input / cached / output, per 1M tokens | Per request | Per month |
|---|---|---|---|
| Claude Haiku 5.5CheapestAnthropic$0.10 in / $0.01 cached / $0.50 out, per 1M tokens | $0.10 / $0.01 / $0.50 | $0.00150 | $22.50$0.00150 a request |
| GPT-6 LunaCheapestOpenAI$0.10 in / $0.01 cached / $0.50 out, per 1M tokens | $0.10 / $0.01 / $0.50 | $0.00150 | $22.50$0.00150 a request |
| Gemini 3.5 Flash-LiteGoogle$0.30 in / $0.03 cached / $2.50 out, per 1M tokens | $0.30 / $0.03 / $2.50 | $0.00550 | $82.50$0.00550 a request |
| Gemini 3.8 FlashGoogleIntroductory price through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached input, $7.50 output.$0.75 in / $0.075 cached / $3.75 out, per 1M tokens | $0.75 / $0.075 / $3.75 | $0.0113 | $169$0.0113 a request |
| Claude Sonnet 5.5Anthropic$2 in / $0.10 cached / $10 out, per 1M tokens | $2 / $0.10 / $10 | $0.0300 | $450$0.0300 a request |
| GPT-6.1 SolOpenAI$2 in / $0.10 cached / $10 out, per 1M tokens | $2 / $0.10 / $10 | $0.0300 | $450$0.0300 a request |
| Gemini 3.1 Pro PreviewGoogle$2 in / $0.20 cached / $12 out, per 1M tokens | $2 / $0.20 / $12 | $0.0320 | $480$0.0320 a request |
| Claude Opus 5.5Anthropic$4 in / $0.20 cached / $20 out, per 1M tokens | $4 / $0.20 / $20 | $0.0600 | $900$0.0600 a request |
| Claude Fable 5.1Anthropic$10 in / $0.25 cached / $50 out, per 1M tokens | $10 / $0.25 / $50 | $0.1500 | $2,250$0.1500 a request |
| GPT-6 AstraOpenAI$10 in / $1 cached / $50 out, per 1M tokens | $10 / $1 / $50 | $0.1500 | $2,250$0.1500 a request |
How the cost is calculated
Each request costs its uncached input tokens at the model's input rate, its cached input tokens at the cached-input rate, and its output tokens at the output rate, with every rate quoted per million tokens. The monthly figure is that cost times your requests per day times 30.
Some models charge more once a prompt passes a length threshold, and then bill the whole request at the higher rate. The table switches to that rate when your input tokens pass it and says so on the row.
The figures are standard prices for text. Batch and flex processing, cache writes and storage, regional processing, and tools such as web search are priced separately and are not included.
Choosing a model takes more than price
The cheapest model that does the job well is usually the right one, and only your own tests can tell you which that is. For how to weigh quality, speed, privacy and deployment alongside cost, see which LLM to use for each kind of work. Model fees are also only one layer of what an agent costs: setup, the monthly service and your team's time are covered in what an AI agent costs.
Frequently asked questions
How is the cost of an LLM API request calculated?
Vendors charge per token, quoted per million tokens, with separate rates for input and output. A request costs its input tokens times the input rate plus its output tokens times the output rate. Input read from a prompt cache is billed at a lower cached-input rate. Output is the expensive side: on every model in this calculator it costs at least five times as much per token as input.
Why can the same text cost different amounts on different models?
Each vendor counts tokens with its own tokenizer, so one prompt is a different number of tokens on each. Anthropic says the tokenizer in Claude 4.7 and later models produces about 30% more tokens for the same text than its earlier one. For a close comparison, measure your own prompts on each model you are considering.
Does prompt caching lower the cost?
Yes, for the part of a prompt that repeats, such as a long system prompt or tool definitions. Cached input is billed at a fraction of the normal input rate. Writing to the cache can cost extra (Anthropic and OpenAI charge more for cache writes, and Google charges for cache storage by the hour), which this calculator leaves out.
Are batch discounts included?
No. The calculator uses standard prices. OpenAI and Google both list batch prices at half their standard rates for work that can wait, so a workload that does not need an answer right away can cost much less than shown here.
How current are the prices?
Every price comes from the vendor's official pricing page, linked on this page, and the date we last checked is shown above the results. We recheck the pages and update the calculator at least once a month.
Sources
Let’s talk aboutyour business.
We build custom AI agents, and we offer fractional CTO advisory and senior-level custom software development. It starts with a conversation.