Free calculator · AI
Compare LLM API costs and a self-hosting estimate
Enter requests per day and tokens per request, then compare up to four current models from Anthropic, OpenAI and Google, plus a custom price. It shows monthly, per-1,000-request and yearly cost, the effect of caching and batch, and what the same volume would cost on rented GPUs.
How this is calculated
The calculator prices one request on each chosen model from its standard per-million-token prices, applies the share of input served from the prompt cache and the share of requests sent through the Batch API, and multiplies by your monthly requests. For self-hosting, it sizes GPUs for peak token throughput and adds a monthly overhead.
Step by step
- Requests per month = requests per day × days per month (30.4 for every day, 22 for business days).
- Cost per request = (input tokens × uncached share × input price + input tokens × cached share × cache-hit price + output tokens × output price) ÷ 1,000,000. A model without a cache-hit price uses its input price.
- Batch: cost × batch share × (1 − batch discount) + cost × (1 − batch share).
- Monthly cost = cost per request × requests per month; per 1,000 requests and per year follow from it.
- Self-hosting GPUs = ceiling(daily tokens ÷ 86,400 × peak-to-average ÷ (tokens per second per GPU × utilisation)), at least one when there is traffic.
- Self-hosting per month = GPUs × GPU price per hour × 730 hours + monthly overhead.
- Break-even: the first requests per day at which the self-hosting cost is no higher than the model's API cost, or none when one GPU's worth of traffic costs less on the API than the GPU itself.
Default assumptions
Assumptions marked adjustable can be changed in the calculator; the others are fixed parts of the model.
| Assumption | Default | Sources |
|---|---|---|
| Requests per day (example)adjustable | 50,000 | None |
| Input tokens per request (example)adjustable | 1,500 tokens | None |
| Output tokens per request (example)adjustable | 400 tokens | None |
| Days of traffic per monthadjustable | 30.4 days | None |
| Custom row: input priceadjustable | $1 per million tokens | None |
| Custom row: output priceadjustable | $4 per million tokens | None |
| Custom row: cached input priceadjustable | $0.10 per million tokens | |
| Custom row: batch discountadjustable | 50% | |
| Share of input tokens served from the prompt cacheadjustable | 0% | None |
| Share of requests sent through the Batch APIadjustable | 0% | None |
| Self-hosting: GPU price per houradjustable | $4 per hour | |
| Self-hosting: sustained tokens per second per GPUadjustable | 1,500 tokens/s | |
| Self-hosting: share of that throughput used at peakadjustable | 50% | None |
| Self-hosting: peak traffic ÷ average trafficadjustable | 3.0× | None |
| Self-hosting: monthly running overheadadjustable | $3,000 a month | None |
| Hours an always-on GPU runs per month | 730 hours | |
| Anthropic Claude Opus 5.5: input / cached input / output per million tokens | $4 / $0.20 / $20 (checked 9 Oct 2026) | |
| Anthropic Claude Sonnet 5.5: input / cached input / output per million tokens | $2 / $0.10 / $10 (checked 9 Oct 2026) | |
| Anthropic Claude Haiku 5.5: input / cached input / output per million tokens | $0.10 / $0.01 / $0.50 (checked 9 Oct 2026) | |
| OpenAI GPT-6 Astra: input / cached input / output per million tokens | $10 / $1 / $50 (checked 9 Oct 2026) | |
| OpenAI GPT-6.1 Sol: input / cached input / output per million tokens | $2 / $0.10 / $10 (checked 9 Oct 2026) | |
| OpenAI GPT-6 Luna: input / cached input / output per million tokens | $0.10 / $0.01 / $0.50 (checked 9 Oct 2026) | |
| Google Gemini 3.1 Pro (preview): input / cached input / output per million tokens | $2 / $0.20 / $12 (checked 9 Oct 2026) | |
| Google Gemini 3.8 Flash: input / cached input / output per million tokens | $0.75 / $0.08 / $3.75 (checked 9 Oct 2026) | |
| Google Gemini 3.5 Flash-Lite: input / cached input / output per million tokens | $0.30 / $0.03 / $2.50 (checked 9 Oct 2026) |
What this doesn’t model
- Prices are the providers' standard pay-as-you-go list prices on the date shown; they change often, so check the linked pages before you budget.
- Token counts are averages you enter. Each provider counts tokens with its own tokenizer, so the same text can be a different number of tokens on each model.
- The self-hosting estimate covers rented GPUs and an overhead you enter. It assumes your throughput figure holds at your prompt lengths and that traffic fits the peak-to-average ratio.
- A self-hosted open-weight model is not quality-equivalent to the hosted models; the comparison is of cost only.
- Higher rates for very long prompts (over 100K tokens on Claude Haiku 5.5, 200K on Gemini 3.1 Pro, 272K on the OpenAI models); the tool flags when your input is above them.
- Prompt-cache write premiums (1.25× or 2× the input price on Claude) and cache storage charges (per million tokens per hour on Gemini).
- Priority, fast and flex tiers, regional or data-residency premiums, and prices on other clouds such as Bedrock or Vertex AI.
- Charges for server-side tools such as web search, and tokens added by tool definitions.
- Taxes, credits, committed-use and other negotiated discounts.
- Differences in answer quality, speed and tokenizer between models; the same text can be a different number of tokens on each.
Sources
- Anthropic, Pricing (Claude API documentation). Accessed . Per-MTok prices for input, cache hits and output; the Batch API is 50% off input and output; cache writes cost 1.25× (5-minute) or 2× (1-hour) the input price.
- OpenAI, Pricing (OpenAI API documentation). Accessed . Standard and Batch tables per 1M tokens; every Batch price listed for these models is half the Standard price. Long-context rates apply above 272K input tokens.
- Google, Gemini Developer API pricing. Accessed . Paid-tier Standard prices per 1M tokens; Batch API is a 50% cost reduction. Gemini 3.1 Pro has a higher rate for prompts over 200k tokens.
- Lambda, GPU cloud pricing. Accessed . On-demand NVIDIA H100 SXM (80 GB) from $3.99 per GPU-hour in 8-GPU instances to $4.29 for a single GPU, plus applicable tax.
- vLLM project, Benchmark CLI (vLLM documentation). Accessed . `vllm bench serve` reports request throughput and total token throughput (tok/s) for your model, hardware and prompt lengths.
- Amazon Web Services, AWS Pricing Calculator assumptions. Accessed . The calculator assumes 730 hours in a month (365 days × 24 hours ÷ 12 months).
Last reviewed by the QuantmHill engineering team. Found an error?
Link to or cite this tool
Writing about this topic? Link to the calculator or cite it. Its method, defaults and sources are all on this page, so readers can check the numbers.
Embed this calculator
You can put this calculator on your own site for free. Paste the code below where it should appear. It loads the same calculator in a frame, with a link back to this page for the full method and sources.
The credit line links to this page with the anchor text “QuantmHill”. You may edit it, add rel="nofollow" or remove it — the calculator works the same either way. Add ?theme=light or ?theme=dark to the iframe address to fix its colour scheme; otherwise it follows the visitor's system setting.
Add this once per page, after the iframe, if you want the frame to grow and shrink with the calculator instead of using the fixed height above. It accepts messages from quantmhill.com only and resizes only the frame that sent them.
Frequently asked questions
The monthly cost of your traffic on each model you choose: input and output tokens at the provider's standard list price, less any share served from the prompt cache or sent through the Batch API. It also sizes rented GPUs for the same token volume at peak and shows where self-hosting would cost the same as each model.
From the Anthropic, OpenAI and Google pricing pages, read on 9 Oct 2026; each price row lists its source and the date it was checked. Providers change prices and release models often, so treat the table as a dated snapshot and use the custom price for anything newer.
Cached input is billed at the model's cache-hit price, which for the models in the table is 5% to 10% of the normal input price. All three providers' pricing pages put the Batch API at half the standard rate for these models, in return for results that can take hours. The tool applies the cache first and the batch discount to what is left.
Only at high, steady volume. GPUs are paid by the hour whether or not they are busy, and they have to be sized for peak traffic, so the tool shows the request volume at which your GPU price, throughput and overhead first match each model's API cost. With the default assumptions, none of the default models reaches that point. The comparison is of cost only: an open-weight model won't answer like the hosted ones.
Cost scales with every token. Long system prompts and retrieved documents add input tokens to every request, and output tokens cost five to about eight times as much as input for the models in the table. On self-hosted GPUs, more tokens per request means fewer requests per GPU, so the GPU count rises too.
Higher rates for very long prompts (over 100K tokens on Claude Haiku 5.5, 200K on Gemini 3.1 Pro, 272K on the OpenAI models); the tool flags when your input is above them. Prompt-cache write premiums (1.25× or 2× the input price on Claude) and cache storage charges (per million tokens per hour on Gemini). Priority, fast and flex tiers, regional or data-residency premiums, and prices on other clouds such as Bedrock or Vertex AI. Charges for server-side tools such as web search, and tokens added by tool definitions. Taxes, credits, committed-use and other negotiated discounts. Differences in answer quality, speed and tokenizer between models; the same text can be a different number of tokens on each.
Want an engineer to check your numbers?
Send us your inputs and the decision you're weighing. We'll reply within one business day with an honest read on whether we can help.