Free · No sign-up · Runs in your browser

LLM API Cost Calculator

Estimate what an AI feature will cost before you build it. Enter your tokens per request and traffic to compare Claude, GPT and Gemini models side by side.

Your usage

Prompt, instructions and context
Length of the response

Providers

Paste a typical prompt or document to estimate its tokens, then apply it as input or output.

≈ 0 tokens

Monthly cost by model

30,000 requests / month

  • GPT-6 Luna cheapest

    $0.10 in · $0.50 out per 1M

    $12.00

    per month

    $0.00040 / request$0.4000 / day

  • Gemini 3.5 Flash-Lite

    $0.30 in · $2.50 out per 1M

    $51.00

    per month

    $0.00170 / request$1.70 / day

  • Gemini 3.8 Flash

    $0.75 in · $3.75 out per 1M

    $90.00

    per month

    $0.00300 / request$3.00 / day

    Introductory price until Dec 31, 2026 ($1.50 / $7.50 after)

  • Claude Haiku 4.5

    $1.00 in · $5.00 out per 1M

    $120.00

    per month

    $0.00400 / request$4.00 / day

  • Claude Sonnet 5.5

    $2.00 in · $10.00 out per 1M

    $240.00

    per month

    $0.00800 / request$8.00 / day

  • GPT-6.1 Sol

    $2.00 in · $10.00 out per 1M

    $240.00

    per month

    $0.00800 / request$8.00 / day

  • Gemini 3.1 Pro Preview

    $2.00 in · $12.00 out per 1M

    $270.00

    per month

    $0.00900 / request$9.00 / day

    Higher price for prompts over 200k tokens

  • Claude Opus 5.5

    $4.00 in · $20.00 out per 1M

    $480.00

    per month

    $0.0160 / request$16.00 / day

  • Claude Fable 5.1

    $10.00 in · $50.00 out per 1M

    $1,200

    per month

    $0.0400 / request$40.00 / day

  • GPT-6 Astra

    $10.00 in · $50.00 out per 1M

    $1,200

    per month

    $0.0400 / request$40.00 / day

Standard list prices in USD, last checked October 6, 2026 from the official pricing pages of Anthropic, OpenAI and Google. Excludes caching and batch discounts, taxes and other fees. Token counts are estimates.

How LLM API pricing works

Large language model APIs charge by the token. Every request has input tokens (your instructions, the user's message and any context you send) and output tokens (the model's response). Each is priced per million tokens, and output is usually several times more expensive than input.

The cost of one request is:

(input tokens × input price + output tokens × output price) ÷ 1,000,000

Multiply by your requests per day and days per month to get the monthly bill. The calculator does this for every model at once, so you can see the trade-off between capability and cost.

How to estimate your tokens

  • One token is roughly 4 characters or three quarters of an English word.
  • A system prompt of 300 words is around 400 tokens and is sent with every request.
  • Retrieved documents (for example in RAG) often make up most of the input tokens.
  • Chat apps resend the conversation history, so input grows with every turn.

Paste a real prompt into the estimator above for a quick count. For exact numbers, use the token counting tool your provider offers before launch.

How to choose the right model

The cheapest model is not always the cheapest system. A stronger model that gets the answer right first time can cost less than a smaller one that needs retries or human review. A practical approach is to start with a capable model, measure quality on real examples, then test smaller models and keep the cheapest one that still meets your quality bar.

Five ways to reduce LLM costs

  1. Send less context. Retrieve only the relevant passages instead of whole documents. The RAG architecture blueprint shows how.
  2. Cache repeated prompts. Long system prompts and shared documents can be cached at a much lower price on the major providers.
  3. Keep outputs short. Ask for structured, concise answers, since output tokens cost the most.
  4. Batch work that can wait. Batch APIs process non-urgent jobs at a discount.
  5. Route by difficulty. Send simple requests to a small model and only hard ones to a large model.

Planning an AI feature?

I build AI features and LLM integrations with production backends, from Claude-powered product listings to automated content pipelines. See my AI workflow automation services for how I keep AI features reliable and cost-efficient.

Frequently asked questions

How is LLM API cost calculated?

Providers charge separately for input tokens (your prompt and context) and output tokens (the response), priced per million tokens. The cost of one request is input tokens times the input price plus output tokens times the output price, divided by one million.

How many tokens is a word?

In English, one token is roughly 4 characters or about three quarters of a word, so 1,000 words is around 1,300 tokens. Exact counts depend on each model's tokenizer, so this calculator gives an estimate.

Why are output tokens more expensive?

Generating output takes more computation than reading input, so most providers price output tokens several times higher. Keeping responses concise is one of the easiest ways to cut cost.

How can I reduce my LLM API costs?

Use the smallest model that meets your quality bar, cache repeated context such as long system prompts, send only the relevant context (for example with retrieval), keep outputs short and use batch processing for work that is not time sensitive.

Want this built for you?

I build custom automations, AI features and production backends. Describe your project on Upwork and get a clear plan, timeline and cost within 2 hours.

Hire Me on Upwork