Guides

What Are Reasoning Models? And When They Are Worth the Cost

Reasoning models think before answering, which helps on hard problems and wastes money on easy ones. How they work, how they bill, and how to use them well.

5 min read Clean APIs Team
What Are Reasoning Models? And When They Are Worth the Cost
Contents

Reasoning models are the biggest shift in AI coding assistance since tool calling — and the easiest to use badly. They solve problems ordinary models get wrong, and they quietly cost several times more per task.

This explains what they actually do, how billing works, and where they earn their price.

What a reasoning model does differently#

An ordinary model predicts its answer directly. A reasoning model first generates an internal chain of thought — working through the problem, considering approaches, checking itself — and only then produces the answer.

Concretely, ask "a train covers 120 km in 1.5 hours, what is its speed in m/s?":

Ordinary model: produces an answer immediately. Often right, sometimes confidently wrong on the unit conversion.

Reasoning model: internally works through 120 km → 120,000 m, 1.5 h → 5,400 s, 120,000 / 5,400 = 22.22, then answers.

That intermediate work is real generated text. You do not have to display it, but it exists — and it is billed.

How the response is shaped#

Reasoning arrives in a separate field so it never contaminates your answer:

{
  "choices": [{
    "message": {
      "role": "assistant",
      "content": "The speed is 22.22 m/s.",
      "reasoning_content": "120 km in 1.5 hours. 120 km = 120000 m. 1.5 h = 5400 s. 120000 / 5400 = 22.22 m/s."
    },
    "finish_reason": "stop"
  }]
}

content is what you show the user. reasoning_content is optional — put it in a collapsible panel, log it for debugging, or ignore it entirely.

message = response.choices[0].message

print("Answer:", message.content)

reasoning = getattr(message, "reasoning_content", None)
if reasoning:
    print("\nThinking:", reasoning)

While streaming, reasoning arrives as delta.reasoning_content and usually finishes before the first delta.content chunk — so you can show a "thinking…" state and then swap to the answer.

The billing reality#

Reasoning tokens count as completion tokens. This is the single most important thing to understand.

A reasoning model answering a trivial question might produce:

  • 300 tokens of internal reasoning
  • 5 tokens of visible answer

You pay for 305. The usage block reflects it:

"usage": {
  "prompt_tokens": 20,
  "completion_tokens": 305,
  "total_tokens": 325
}

The practical consequence: a reasoning model at the same headline rate as a non-reasoning one costs several times more per task, because output volume is several times higher.

The max_tokens trap#

This catches almost everyone once.

If max_tokens is small and the model spends it all on reasoning, content comes back empty. You paid full price and got nothing.

# Broken: 500 tokens may be entirely consumed by thinking
response = client.chat.completions.create(
    model="claude-opus-4.8",
    messages=[{"role": "user", "content": "Find the deadlock in this code..."}],
    max_tokens=500,
)
print(response.choices[0].message.content)   # possibly ""

Clean APIs guards against the worst case: max_tokens below 2048 is raised to 2048, and omitting it defaults to 8192. That means a naive request still returns something useful.

For genuinely hard problems, set it explicitly:

Task difficulty Suggested max_tokens
Short answers 2048
General use 8192 (our default)
Complex debugging, long code output 16000+

Where reasoning genuinely wins#

Based on the kinds of problems where the internal working changes the outcome:

Debugging subtle bugs. Race conditions, off-by-one errors, incorrect assumptions about async ordering. The model catches its own wrong first guess.

Architecture decisions. Trade-offs between approaches, where the reasoning is the value.

Algorithm work. Complexity analysis, correctness arguments, edge case enumeration.

Multi-step planning. Breaking a large task into ordered steps with dependencies.

Anything with maths. Unit conversions, financial calculations, geometry.

Where it is a waste#

Formatting and style. No reasoning required.

Renaming and mechanical refactors. Deterministic transformations.

Docstrings and comments. Describing existing code.

Translation and summarising. Sequence-to-sequence work.

Simple CRUD. Predictable patterns.

For all of these, a fast non-reasoning model gives an identical result for a fraction of the cost — and returns it sooner.

A practical split#

Because every Clean APIs plan reaches all 31 models, you can switch per task with a one-string change:

FAST_MODEL = "some-fast-model"          # formatting, renames, docs
DEEP_MODEL = "claude-opus-4.8"        # debugging, design, algorithms

def ask(prompt: str, hard: bool = False):
    return client.chat.completions.create(
        model=DEEP_MODEL if hard else FAST_MODEL,
        messages=[{"role": "user", "content": prompt}],
        max_tokens=16000 if hard else 4096,
    )

In agent tools the same idea applies at the settings level: keep a cheap model as your default and switch when you hit something genuinely difficult.

Showing reasoning to users#

If you build a product on top of this, three patterns work well:

Collapsed by default. A "Show thinking" toggle. Most users do not want it; the ones debugging do.

Live during generation, hidden after. Show reasoning streaming in as a progress indicator, collapse it when the answer arrives. This is what our playground does.

Logged, never shown. Store it for your own debugging. When a model gets something wrong, the reasoning usually shows exactly where it went off.

What does not work is interleaving it with the answer — users cannot tell which part is the conclusion.

Identifying reasoning models#

Check the reasoning capability. On the models page, filter by the 🧠 Reasoning badge. Programmatically:

curl https://cleanapis.com/v1/models \
  -H "Authorization: Bearer cc_your_key_here"
{
  "id": "claude-opus-4.8",
  "capabilities": ["reasoning", "vision", "tools", "streaming"],
  "supported_parameters": ["reasoning", "include_reasoning", "max_tokens", "tools"]
}

Note that not every reasoning model exposes its chain of thought — some reason internally without returning reasoning_content. You still pay for those tokens; you just cannot read them.

Measuring whether it is worth it#

The honest way to decide is to compare on your own work.

Run the same ten representative tasks through a reasoning and a non-reasoning model. Then look at Dashboard → Usage, which logs the real token count and cost of every request, split by model.

If the reasoning model costs 4× and solves two extra problems out of ten, that may be an easy yes. If it costs 4× and produces the same output, it is an easy no. The data tells you; intuition does not.

Next steps#

Get a free API key — 5M tokens monthly, reasoning models included on every plan.

Ready to build?

Everything in this article works on the free tier — 5M tokens every month, all 33 models, no card.

Related reading