Reasoning Models

How reasoning models think before answering, and how to budget tokens for them.

On this page

Reasoning Models#

Models advertising the Reasoning capability work through a problem internally before producing an answer. This markedly improves maths, logic, multi-step planning, and debugging.

Reasoning content#

When a model exposes its thinking, it arrives in a separate field so it never contaminates the answer:

{
  "choices": [{
    "message": {
      "role": "assistant",
      "content": "The speed is 22.22 m/s.",
      "reasoning_content": "120 km in 1.5 hours. 120 km = 120000 m. 1.5 h = 5400 s. 120000 / 5400 = 22.22 m/s."
    },
    "finish_reason": "stop"
  }]
}

content holds the answer to show your user. reasoning_content is optional — display it in a collapsible panel or ignore it.

While streaming, reasoning arrives as delta.reasoning_content and typically completes before the first delta.content chunk.

Token budget#

This is the single most common mistake with reasoning models.

Reasoning tokens count toward completion_tokens. A tight max_tokens can be spent entirely on reasoning, leaving an empty content.

We guard against this: max_tokens below 2048 is raised to 2048, and omitting it defaults to 8192. For hard problems, allow 16000 or more.

Task Suggested max_tokens
Short answers 2048
General use 8192 (default)
Complex reasoning, long code 16000+

Choosing a model#

Reasoning models are slower and cost more, since you pay for reasoning tokens too. Prefer them for algorithm design, debugging, maths, and multi-step planning; prefer non-reasoning models for summarising, formatting, translation, and simple lookups.

Example#

from openai import OpenAI

client = OpenAI(api_key="cc_your_key_here", base_url="https://cleanapis.com/v1")

response = client.chat.completions.create(
    model="claude-opus-4.8",
    messages=[{
        "role": "user",
        "content": "A train covers 120 km in 1.5 hours. What is its speed in m/s?",
    }],
    max_tokens=8192,
)

message = response.choices[0].message
print("Answer:", message.content)

reasoning = getattr(message, "reasoning_content", None)
if reasoning:
    print("\nThinking:", reasoning)

Still stuck?

Open a support ticket from your dashboard and we'll take a look.

Get help