Understanding Token Pricing — What You Actually Pay For
Tokens, context windows, reasoning overhead, and why an agent costs more than a chatbot. A practical guide to estimating and controlling AI API spend.
Contents
AI API pricing confuses people because the unit being sold — a token — is invisible. You cannot count tokens by looking at your prompt, and the same task can cost wildly different amounts depending on which model runs it and how it is configured.
This guide explains what you are actually paying for, how to estimate a bill before you run anything, and where costs hide.
What is a token#
A token is roughly a piece of a word. Not a character, not a whole word — something in between, produced by the model's tokenizer.
Practical rules of thumb for English:
| Input | Approximate tokens |
|---|---|
| 1 token | ~4 characters |
| 1 token | ~0.75 words |
| 100 words | ~130 tokens |
| 1 page of prose (~500 words) | ~650 tokens |
Code is denser. Punctuation, brackets, and identifiers fragment more, so code runs closer to 3 characters per token. A 200-line file is often 2,000–3,000 tokens.
Non-Latin scripts cost more. Bengali, Chinese, Japanese, Korean, and Arabic often consume close to one token per character — the same sentence can cost several times more than its English equivalent.
You pay for input and output#
Every request bills two directions:
- Prompt tokens — everything you send: system prompt, conversation history, file contents, tool definitions
- Completion tokens — everything the model generates
Output usually costs more per token than input, often 3–4×, because generation is computationally heavier than reading.
This is why long conversations get expensive in a non-obvious way: the API is stateless, so every turn resends the entire history.
Turn 1: 500 prompt + 200 completion = 700 tokens
Turn 2: 700 prompt + 200 completion = 900 tokens
Turn 3: 900 prompt + 200 completion = 1,100 tokens
Turn 10: 2,300 prompt + 200 completion = 2,500 tokens
Ten turns of a short conversation is not 7,000 tokens. It is closer to 16,000 — the history is paid for again on every single turn.
Context window is a limit, not a price#
The context window is the maximum combined prompt + completion the model accepts. A 128K window does not mean you pay for 128K — you pay for what you actually send.
But it does set a ceiling on what is possible. Agents send whole files plus their own reasoning history, so a small window fails fast on real work:
| Context | Realistic use |
|---|---|
| 8K | Single short file, simple Q&A |
| 32K | A few files, moderate conversation |
| 128K | Real codebase work, agentic sessions |
| 1M | Entire repositories, very long documents |
Every Clean APIs model lists its context window on the models page, and returns it as context_length from GET /v1/models.
Reasoning models: the hidden multiplier#
This is where most unexpected bills come from.
Reasoning models think before answering. That thinking is generated text, and it counts as completion tokens — even though you may never display it.
A reasoning model answering "what is 2+2" might produce:
- 300 tokens of internal reasoning
- 5 tokens of visible answer
You pay for 305.
Two consequences:
Costs are higher than they look. A reasoning model at the same headline rate as a non-reasoning one will cost several times more per task, because it generates far more.
Small max_tokens breaks them. If you cap output at 500 and the model spends 500 on reasoning, content comes back empty. You paid, and got nothing useful.
Clean APIs guards against the second problem: max_tokens below 2048 is raised to 2048, and omitting it defaults to 8192. For genuinely hard problems, set 16000 or more.
response = client.chat.completions.create(
model="claude-opus-4.8",
messages=[{"role": "user", "content": "Debug this race condition..."}],
max_tokens=16000, # room to think AND answer
)
message = response.choices[0].message
print("Answer:", message.content)
# Reasoning arrives separately — display it or ignore it
reasoning = getattr(message, "reasoning_content", None)
Tool definitions cost tokens too#
If you send tools, those JSON schemas are part of the prompt on every request of the conversation.
Five tools with detailed descriptions and parameter schemas can easily be 1,500 tokens. In a twenty-turn agent session, that is 30,000 tokens spent purely on repeating definitions the model has already seen.
Practical response: keep tool descriptions tight, and only send the tools relevant to the current task rather than your entire catalogue.
Images cost tokens#
Vision input is billed as prompt tokens based on resolution. A typical screenshot runs a few hundred to a few thousand tokens. Higher resolution costs more, so downscale before sending when detail is not needed.
How Clean APIs bills#
Two budgets, spent in order:
1. Monthly plan allowance. Every plan includes a token allowance that resets monthly. Requests draw from it first, at no per-request cost.
| Plan | Tokens / month | Price | Per 1M |
|---|---|---|---|
| Free | 5M | $0 | — |
| Starter | 50M | $3.50 | $0.070 |
| Basic | 100M | $5.75 | $0.0575 |
| Pro | 200M | $9.50 | $0.0475 |
| Scale | 500M | $15.50 | $0.031 |
| Unlimited | No cap | $50 | — |
2. Account balance. Once the allowance is spent, requests bill against your balance at each model's published per-token rate.
If both are empty, requests return 402 before reaching a model. Nothing is charged for a rejected request, and you are never silently cut off mid-response.
Larger packages cost less per token — Scale is roughly 42% cheaper per million than Starter. Every plan reaches every model; you are buying volume, never access.
Estimating a bill before you spend#
Work out your average request, then multiply.
A chatbot turn:
System prompt: 200 tokens
History (5 turns): 2,000 tokens
User message: 100 tokens
────────────────────────────────
Prompt: 2,300 tokens
Completion: 300 tokens
Total: 2,600 tokens per turn
At 200M tokens/month (Pro, $5.75) that is roughly 77,000 turns — about 2,500 per day.
An agent task (file edit):
System + tools: 1,500 tokens
File contents: 3,000 tokens
Reasoning: 2,000 tokens
Tool calls + result:1,000 tokens
Final answer: 500 tokens
────────────────────────────────
Total: 8,000 tokens per task
Same 200M allowance: about 25,000 agent tasks, or 800 a day.
Agents cost roughly 3× a chatbot turn per unit of work, and that is before multi-turn loops where the agent reads, edits, tests, and re-reads.
Six ways to cut spend#
Trim conversation history. Keep the system prompt plus the last N turns instead of everything. This is the single largest saving in most chat applications.
Match the model to the task. Reserve reasoning models for problems that need reasoning. Formatting, summarising, translation, and classification do not.
Cap max_tokens deliberately. Not too low (reasoning models return nothing), not unlimited. Set it to what a good answer actually needs.
Send fewer tools. Only the ones relevant right now.
Downscale images. Send the resolution the task requires, not whatever the screenshot happened to be.
Watch the real numbers. Clean APIs's Dashboard → Usage logs every request with its exact token count and cost to six decimal places — per key and per model. Estimates are useful for planning; real data tells you where the money actually went.
Set a usage alert in Settings → Notifications to get warned at whatever percentage of your allowance you choose.
Reading the usage response#
Every completion returns exactly what it cost:
{
"usage": {
"prompt_tokens": 2300,
"completion_tokens": 300,
"total_tokens": 2600
}
}
For streamed requests, the usage block arrives in the final chunk before data: [DONE]. Log it — you cannot optimise what you do not measure.
Next steps#
- Compare model prices — filter by capability and cost per 1M
- Pricing plans — full comparison
- Chat completions reference — every parameter
Start free with 5M tokens a month and watch your real numbers before committing to anything.
Ready to build?
Everything in this article works on the free tier — 5M tokens every month, all 33 models, no card.