Skip to main content
VantaSoft
Resources

LLM Cost Calculator.

Enter the tokens in a typical request and how many you send a day. You'll see what each current Claude, GPT and Gemini model would cost, cheapest first.

Your workload

tokens

Everything you send: system prompt, tool definitions, history and the new message.

tokens

What the model writes back, including any reasoning tokens it bills.

requests

API calls, not conversations. An agent often makes several per task.

%

The repeated part of your prompts, if you use prompt caching.

Cost by model

Standard API prices, checked October 8, 2026. A month is 30 days.

Monthly and per-request cost for each model, cheapest first
ModelPer month
Claude Haiku 5.5CheapestAnthropic$0.10 in / $0.01 cached / $0.50 out, per 1M tokens$22.50$0.00150 a request
GPT-6 LunaCheapestOpenAI$0.10 in / $0.01 cached / $0.50 out, per 1M tokens$22.50$0.00150 a request
Gemini 3.5 Flash-LiteGoogle$0.30 in / $0.03 cached / $2.50 out, per 1M tokens$82.50$0.00550 a request
Gemini 3.8 FlashGoogleIntroductory price through December 31, 2026. From January 1, 2027: $1.50 input, $0.15 cached input, $7.50 output.$0.75 in / $0.075 cached / $3.75 out, per 1M tokens$169$0.0113 a request
Claude Sonnet 5.5Anthropic$2 in / $0.10 cached / $10 out, per 1M tokens$450$0.0300 a request
GPT-6.1 SolOpenAI$2 in / $0.10 cached / $10 out, per 1M tokens$450$0.0300 a request
Gemini 3.1 Pro PreviewGoogle$2 in / $0.20 cached / $12 out, per 1M tokens$480$0.0320 a request
Claude Opus 5.5Anthropic$4 in / $0.20 cached / $20 out, per 1M tokens$900$0.0600 a request
Claude Fable 5.1Anthropic$10 in / $0.25 cached / $50 out, per 1M tokens$2,250$0.1500 a request
GPT-6 AstraOpenAI$10 in / $1 cached / $50 out, per 1M tokens$2,250$0.1500 a request

How the cost is calculated

Each request costs its uncached input tokens at the model's input rate, its cached input tokens at the cached-input rate, and its output tokens at the output rate, with every rate quoted per million tokens. The monthly figure is that cost times your requests per day times 30.

Some models charge more once a prompt passes a length threshold, and then bill the whole request at the higher rate. The table switches to that rate when your input tokens pass it and says so on the row.

The figures are standard prices for text. Batch and flex processing, cache writes and storage, regional processing, and tools such as web search are priced separately and are not included.

Choosing a model takes more than price

The cheapest model that does the job well is usually the right one, and only your own tests can tell you which that is. For how to weigh quality, speed, privacy and deployment alongside cost, see which LLM to use for each kind of work. Model fees are also only one layer of what an agent costs: setup, the monthly service and your team's time are covered in what an AI agent costs.

Frequently asked questions

How is the cost of an LLM API request calculated?

Vendors charge per token, quoted per million tokens, with separate rates for input and output. A request costs its input tokens times the input rate plus its output tokens times the output rate. Input read from a prompt cache is billed at a lower cached-input rate. Output is the expensive side: on every model in this calculator it costs at least five times as much per token as input.

Why can the same text cost different amounts on different models?

Each vendor counts tokens with its own tokenizer, so one prompt is a different number of tokens on each. Anthropic says the tokenizer in Claude 4.7 and later models produces about 30% more tokens for the same text than its earlier one. For a close comparison, measure your own prompts on each model you are considering.

Does prompt caching lower the cost?

Yes, for the part of a prompt that repeats, such as a long system prompt or tool definitions. Cached input is billed at a fraction of the normal input rate. Writing to the cache can cost extra (Anthropic and OpenAI charge more for cache writes, and Google charges for cache storage by the hour), which this calculator leaves out.

Are batch discounts included?

No. The calculator uses standard prices. OpenAI and Google both list batch prices at half their standard rates for work that can wait, so a workload that does not need an answer right away can cost much less than shown here.

How current are the prices?

Every price comes from the vendor's official pricing page, linked on this page, and the date we last checked is shown above the results. We recheck the pages and update the calculator at least once a month.

Sources

Let’s talk aboutyour business.

We build custom AI agents, and we offer fractional CTO advisory and senior-level custom software development. It starts with a conversation.