How LLM API pricing works
Large language model APIs charge by the token. Every request has input tokens (your instructions, the user's message and any context you send) and output tokens (the model's response). Each is priced per million tokens, and output is usually several times more expensive than input.
The cost of one request is:
(input tokens × input price + output tokens × output price) ÷ 1,000,000
Multiply by your requests per day and days per month to get the monthly bill. The calculator does this for every model at once, so you can see the trade-off between capability and cost.
How to estimate your tokens
- One token is roughly 4 characters or three quarters of an English word.
- A system prompt of 300 words is around 400 tokens and is sent with every request.
- Retrieved documents (for example in RAG) often make up most of the input tokens.
- Chat apps resend the conversation history, so input grows with every turn.
Paste a real prompt into the estimator above for a quick count. For exact numbers, use the token counting tool your provider offers before launch.
How to choose the right model
The cheapest model is not always the cheapest system. A stronger model that gets the answer right first time can cost less than a smaller one that needs retries or human review. A practical approach is to start with a capable model, measure quality on real examples, then test smaller models and keep the cheapest one that still meets your quality bar.
Five ways to reduce LLM costs
- Send less context. Retrieve only the relevant passages instead of whole documents. The RAG architecture blueprint shows how.
- Cache repeated prompts. Long system prompts and shared documents can be cached at a much lower price on the major providers.
- Keep outputs short. Ask for structured, concise answers, since output tokens cost the most.
- Batch work that can wait. Batch APIs process non-urgent jobs at a discount.
- Route by difficulty. Send simple requests to a small model and only hard ones to a large model.
Planning an AI feature?
I build AI features and LLM integrations with production backends, from Claude-powered product listings to automated content pipelines. See my AI workflow automation services for how I keep AI features reliable and cost-efficient.