The Bill Always Arrives After the Decision
The usual order of operations is backwards: ship a feature that calls an LLM API, watch usage ramp up, then open the bill and find out which part of the prompt was expensive. Estimating cost before you ship means understanding three things: how tokens are counted, why input and output are priced so differently, and which few levers actually move the number.
Step 1: A Token Isn't a Word
Every major provider bills by token, not by character or word. A token is a sub-word unit produced by the model's tokenizer — common short words are usually a single token, but longer words, code, and especially non-English text often split into several.
"The cat sat." → 4 tokens
"Antidisestablishmentarianism" → 6+ tokens (one unusual word, several pieces)
"const x = {a: 1};" → 9+ tokens (punctuation and symbols each cost something)
The commonly cited rule of thumb for English prose is about 4 characters per token, or roughly 0.75 tokens per word — useful for a fast mental estimate, but it drifts further from reality the more a prompt consists of code, JSON, or non-English text, since those tokenize less predictably.
Step 2: Input and Output Are Priced Completely Differently
This is the part that catches people off guard: output tokens cost 4 to 6 times more than input tokens, across every major provider.
| Model | Input $/1M tokens | Output $/1M tokens | Output multiple |
|---|---|---|---|
| Claude Opus 5.5 | $4.00 | $20.00 | 5x |
| Claude Sonnet 5.5 | $2.00 | $10.00 | 5x |
| GPT-5 | $1.25 | $10.00 | 8x |
| Gemini 2.5 Flash | $0.30 | $2.50 | ~8.3x |
Why: processing a long input prompt can be parallelized across the whole sequence at once. Generating output can't — each token depends on every token generated before it, so the model runs one full forward pass per output token, sequentially. That's strictly more compute per token, and the price reflects it.
The practical consequence: a prompt that asks for a long, detailed response costs far more than the same prompt asking for a short one — even if the input is identical. If you're building something that can get away with a terser response format (structured output, a capped length, a summary instead of a full rewrite), that's often the single cheapest change available.
Step 3: Estimate Before You Build, Not After
Three numbers multiply into your actual cost:
cost = (input_tokens / 1,000,000 × input_price)
+ (output_tokens / 1,000,000 × output_price)
For a single request, this is simple arithmetic once you know roughly how many tokens your prompt and expected response are. For a feature that'll run thousands of times — a background job, a per-user action, an agent loop — multiply by expected call volume before shipping, not after the first invoice.
Two numbers people chronically underestimate:
- System prompt size. A detailed system prompt with examples and instructions gets resent on every single request unless it's cached (see below). A 2,000-token system prompt across 10,000 daily requests is 20 million input tokens a day from the system prompt alone — before the actual user content.
- Conversation history in a multi-turn chat. Each turn in a stateless chat API resends the entire prior conversation. A 20-turn conversation doesn't cost 20x a single turn — it costs roughly 1+2+3+...+20 turns' worth of resent history, since every earlier message gets billed again on every later request.
The One Free Lever: Prompt Caching
If a large, static chunk of your prompt — a system prompt, a reference document, a tool definition list — gets reused across many requests, prompt caching can cut that portion's cost dramatically (commonly cited around 90% cheaper for a cache hit versus a full-price read). This is the highest-leverage, lowest-effort cost optimization available on every major provider's API, and it requires no change to your actual prompt content — just keeping the cached portion's content byte-identical and placing it before the parts that change.
If you're resending the same long system prompt or document set on every call and haven't set up caching for it, that's usually the first thing worth fixing before looking at anything else.
Quick Reference
| Question | Answer |
|---|---|
| Is input or output more expensive per token? | Output — typically 4-6x the input rate |
| What's the fastest way to estimate token count? | ~4 characters per token for English prose (rougher for code/non-English) |
| What's the biggest easy win for multi-call workloads? | Prompt caching on the static/reused portion of the prompt |
| What silently inflates cost in a chat app? | Resending full conversation history on every turn |
| Should I trust an approximate token count for billing? | Only for ballparking — use the provider's exact token-counting endpoint for anything billing-critical |
Try It
ToolNinja's AI Token Counter & Cost Estimator → estimates token count for any text and shows the cost side by side across current Claude, GPT, and Gemini models, with a slider to model expected output length against input — useful for a quick gut-check before you build, not a replacement for an exact pre-launch calculation.
Sources: