Home/Blog/How to Estimate Your LLM API Costs Before You Get the Bill
🧾
llm api costtoken counterestimate api cost

How to Estimate Your LLM API Costs Before You Get the Bill

A practical walkthrough of how LLM API pricing actually works — tokens, input vs output cost, why a long system prompt costs more than you think — so you can estimate spend before you ship, not after.

October 4, 20265 min readby ToolNinja

The Bill Always Arrives After the Decision

The usual order of operations is backwards: ship a feature that calls an LLM API, watch usage ramp up, then open the bill and find out which part of the prompt was expensive. Estimating cost before you ship means understanding three things: how tokens are counted, why input and output are priced so differently, and which few levers actually move the number.


Step 1: A Token Isn't a Word

Every major provider bills by token, not by character or word. A token is a sub-word unit produced by the model's tokenizer — common short words are usually a single token, but longer words, code, and especially non-English text often split into several.

"The cat sat."        → 4 tokens
"Antidisestablishmentarianism" → 6+ tokens (one unusual word, several pieces)
"const x = {a: 1};"   → 9+ tokens (punctuation and symbols each cost something)

The commonly cited rule of thumb for English prose is about 4 characters per token, or roughly 0.75 tokens per word — useful for a fast mental estimate, but it drifts further from reality the more a prompt consists of code, JSON, or non-English text, since those tokenize less predictably.


Step 2: Input and Output Are Priced Completely Differently

This is the part that catches people off guard: output tokens cost 4 to 6 times more than input tokens, across every major provider.

ModelInput $/1M tokensOutput $/1M tokensOutput multiple
Claude Opus 5.5$4.00$20.005x
Claude Sonnet 5.5$2.00$10.005x
GPT-5$1.25$10.008x
Gemini 2.5 Flash$0.30$2.50~8.3x

Why: processing a long input prompt can be parallelized across the whole sequence at once. Generating output can't — each token depends on every token generated before it, so the model runs one full forward pass per output token, sequentially. That's strictly more compute per token, and the price reflects it.

The practical consequence: a prompt that asks for a long, detailed response costs far more than the same prompt asking for a short one — even if the input is identical. If you're building something that can get away with a terser response format (structured output, a capped length, a summary instead of a full rewrite), that's often the single cheapest change available.


Step 3: Estimate Before You Build, Not After

Three numbers multiply into your actual cost:

cost = (input_tokens / 1,000,000 × input_price)
     + (output_tokens / 1,000,000 × output_price)

For a single request, this is simple arithmetic once you know roughly how many tokens your prompt and expected response are. For a feature that'll run thousands of times — a background job, a per-user action, an agent loop — multiply by expected call volume before shipping, not after the first invoice.

Two numbers people chronically underestimate:

  1. System prompt size. A detailed system prompt with examples and instructions gets resent on every single request unless it's cached (see below). A 2,000-token system prompt across 10,000 daily requests is 20 million input tokens a day from the system prompt alone — before the actual user content.
  2. Conversation history in a multi-turn chat. Each turn in a stateless chat API resends the entire prior conversation. A 20-turn conversation doesn't cost 20x a single turn — it costs roughly 1+2+3+...+20 turns' worth of resent history, since every earlier message gets billed again on every later request.

The One Free Lever: Prompt Caching

If a large, static chunk of your prompt — a system prompt, a reference document, a tool definition list — gets reused across many requests, prompt caching can cut that portion's cost dramatically (commonly cited around 90% cheaper for a cache hit versus a full-price read). This is the highest-leverage, lowest-effort cost optimization available on every major provider's API, and it requires no change to your actual prompt content — just keeping the cached portion's content byte-identical and placing it before the parts that change.

If you're resending the same long system prompt or document set on every call and haven't set up caching for it, that's usually the first thing worth fixing before looking at anything else.


Quick Reference

QuestionAnswer
Is input or output more expensive per token?Output — typically 4-6x the input rate
What's the fastest way to estimate token count?~4 characters per token for English prose (rougher for code/non-English)
What's the biggest easy win for multi-call workloads?Prompt caching on the static/reused portion of the prompt
What silently inflates cost in a chat app?Resending full conversation history on every turn
Should I trust an approximate token count for billing?Only for ballparking — use the provider's exact token-counting endpoint for anything billing-critical

Try It

ToolNinja's AI Token Counter & Cost Estimator → estimates token count for any text and shows the cost side by side across current Claude, GPT, and Gemini models, with a slider to model expected output length against input — useful for a quick gut-check before you build, not a replacement for an exact pre-launch calculation.


Sources:

Share:𝕏 Twitterin LinkedIn

Frequently Asked Questions

Why do output tokens cost more than input tokens?

Generating a token requires a full forward pass through the model, run sequentially — it can't be parallelized across tokens the way processing a long input prompt can be. That computational asymmetry is why every major provider (Anthropic, OpenAI, Google) prices output tokens at roughly 4-6x the input rate.

Is a token the same as a word?

No. A token is a sub-word unit — short, common words are usually one token, but longer or unusual words, code, and non-English text often split into multiple tokens. A commonly cited rule of thumb for English prose is about 4 characters per token, or roughly 0.75 tokens per word, but it varies by content type.

What's the single biggest lever for reducing LLM API cost?

Prompt caching, where the provider supports it. A system prompt or set of reference documents that gets reused across many requests can be cached so repeat requests pay a fraction of the normal input-token rate for that cached portion — often around 90% cheaper for cache hits versus a full-price read every time.

Do I need to count tokens exactly, or is an estimate good enough?

For budgeting and comparing models, a close estimate (within 10-15%) is usually enough to make a decision. For anything billing-critical — like a hard spend cap you can't exceed — use the provider's own exact token-counting endpoint rather than an approximation, since estimates can run slightly high or low depending on the content.

🥷 ToolNinja