Pricing

How to Reduce AI API Costs — 12 Techniques That Actually Work

Practical ways to cut token spend without losing quality, ranked by impact. Measured against real usage data, not guesswork.

6 min read Clean APIs Team
How to Reduce AI API Costs — 12 Techniques That Actually Work
Contents

AI API bills grow in ways that surprise people. A feature that cost $5/month in testing costs $200 in production, and the reason is rarely obvious from the code.

These twelve techniques are ordered by typical impact. The first three usually account for most of the saving.

1. Trim conversation history#

Impact: very high. This is the biggest lever in almost every chat application.

The API is stateless. Every request resends the entire conversation, so cost grows quadratically with turn count:

Turn 1:   500 prompt +  200 completion =   700
Turn 5:  2,300 prompt +  200 completion = 2,500
Turn 10: 4,300 prompt +  200 completion = 4,500
Turn 20: 8,300 prompt +  200 completion = 8,500

Twenty turns of a short conversation is not 14,000 tokens. It is closer to 90,000.

Keep the system prompt plus a sliding window:

def build_messages(system, history, user_msg, keep=6):
    return [
        {"role": "system", "content": system},
        *history[-keep:],                       # last N turns only
        {"role": "user", "content": user_msg},
    ]

For long conversations, summarise older turns instead of dropping them:

if len(history) > 12:
    summary = summarise(history[:-6])           # one cheap call
    history = [{"role": "system", "content": f"Earlier context: {summary}"}, *history[-6:]]

One summarisation call replaces thousands of resent tokens on every subsequent turn.

2. Match the model to the task#

Impact: very high.

Reasoning models generate their thinking as billed output. On a task that needs no reasoning, you pay several times more for an identical result.

Task Model
Formatting, renaming, boilerplate Cheap non-reasoning
Docstrings, comments, summaries Cheap non-reasoning
Translation, classification Cheap non-reasoning
Feature implementation Mid-tier
Debugging, architecture, algorithms Reasoning
FAST = "some-fast-model"
DEEP = "claude-opus-4.8"

def complete(prompt, hard=False):
    return client.chat.completions.create(
        model=DEEP if hard else FAST,
        messages=[{"role": "user", "content": prompt}],
        max_tokens=16000 if hard else 2048,
    )

Because every Clean APIs plan reaches all 31 models, routing by task costs nothing to set up.

→ How to choose an AI model for coding

3. Send less context#

Impact: very high for agents and RAG.

Sending a whole 3,000-token file when the relevant function is 80 lines wastes 90% of the prompt.

For code: extract the relevant function or class rather than the file.

For documents: retrieve the top 3 chunks, not the top 20. More context is not automatically better — irrelevant context also degrades output quality.

For agents: let the model request files via tool calls rather than pre-loading everything. It usually asks for less than you would have sent.

4. Cap max_tokens deliberately#

Impact: high.

Without a cap the model decides how long to be, and models are verbose by default.

max_tokens=150      # a summary
max_tokens=500      # a function
max_tokens=2048     # a detailed explanation
max_tokens=16000    # reasoning on a hard problem

One caveat: with reasoning models a small cap can be consumed entirely by thinking, returning an empty answer. Clean APIs floors it at 2048 and defaults to 8192 to prevent that.

5. Ask for less output#

Impact: high, and easy.

Output usually costs 3–4× input, so prompt wording has direct financial effect.

Instead of Ask for
"Explain this code" "Explain this code in 3 bullet points"
"Fix this function" "Show only the lines that change"
"Review this file" "List issues, one line each, no code"
"Write tests" "Write 3 test cases covering edge cases"

"Show only the diff" instead of "show the fixed file" can cut output by 80% on a large file.

6. Trim tool definitions#

Impact: medium to high for agents.

Tool schemas are part of the prompt on every request. Five verbose definitions can be 1,500 tokens, resent across twenty turns — 30,000 tokens spent repeating what the model already saw.

Keep descriptions to one clear line, and send only the tools relevant to the current task rather than your whole catalogue.

→ Tool calling explained

7. Cache repeated answers#

Impact: high where inputs repeat.

If the same question can arrive twice, the second one should not reach the model:

import hashlib, json

def cache_key(model, messages):
    payload = json.dumps({"model": model, "messages": messages}, sort_keys=True)
    return "ai:" + hashlib.sha256(payload.encode()).hexdigest()

def complete_cached(model, messages, ttl=3600):
    key = cache_key(model, messages)

    if hit := cache.get(key):
        return hit

    result = client.chat.completions.create(model=model, messages=messages)
    cache.set(key, result, ttl)
    return result

Documentation Q&A, classification, and templated generation often see 30%+ hit rates.

8. Batch where the API allows it#

Impact: medium.

Embeddings accept an array, and one request for 100 texts is far cheaper in overhead than 100 requests:

response = client.embeddings.create(
    model="your-embedding-model",
    input=["first text", "second text", "third text"],   # up to 2048
)

Results come back in data with an index matching each input, so ordering is guaranteed.

9. Downscale images#

Impact: medium if you use vision.

Image tokens scale with resolution. A 4K screenshot costs several times a 1024px one, and for reading an error message the extra pixels add nothing.

Resize before sending unless fine detail is genuinely required.

10. Stop generating early#

Impact: medium for interactive tools.

Users interrupt constantly. Cancelling a stream stops generation, so tokens after that point are never produced.

Make cancellation obvious and instant in your UI. Note that tokens already generated are still billed — that is unavoidable, and honest providers log them so you see real spend.

11. Use stop sequences#

Impact: low to medium.

If you only need up to a marker, stop there:

response = client.chat.completions.create(
    model="claude-opus-4.8",
    messages=messages,
    stop=["\n\n", "END"],
)

Useful for structured single-value outputs where the model would otherwise keep explaining.

12. Buy volume#

Impact: up to 42% on the same usage.

Per-token rates fall as package size grows:

Plan Tokens / month Price Per 1M
Free 5M $0 —
Starter 50M $3.50 $0.070
Basic 100M $5.75 $0.0575
Pro 200M $9.50 $0.0475
Scale 500M $15.50 $0.031
Unlimited No cap $50 —

Scale is about 42% cheaper per million than Starter. If you consistently use 400M+, moving up pays for itself immediately.

Every plan reaches every model — you are buying volume, never access.

Measure first, optimise second#

None of this matters if you are optimising the wrong thing. Get the data:

Dashboard → Usage shows, per request:

  • Prompt, completion, and total tokens
  • Exact cost to six decimal places
  • Which model handled it
  • Which API key made it
  • Whether plan allowance or balance paid

Break it down by model to find where the money actually goes. It is often one endpoint or one feature, not everything equally.

Set a usage alert in Settings → Notifications at whatever percentage you choose, so growth is visible before it becomes a surprise.

Every response also returns exactly what it cost:

"usage": {
  "prompt_tokens": 2300,
  "completion_tokens": 300,
  "total_tokens": 2600
}

For streams, that block arrives in the final chunk before [DONE]. Log it — you cannot optimise what you do not measure.

A realistic starting point#

Apply the first three techniques and most applications see a 50–70% reduction without any quality loss:

  1. Sliding-window history
  2. Route easy tasks to a cheap model
  3. Send only relevant context

Then measure again and decide whether the rest is worth the engineering time.

Next steps#

Start free — 5M tokens monthly, and real usage data from your first request.

Ready to build?

Everything in this article works on the free tier — 5M tokens every month, all 33 models, no card.

Related reading