Reasoning Models
How reasoning models think before answering, and how to budget tokens for them.
On this page
Reasoning Models#
Models advertising the Reasoning capability work through a problem internally before producing an answer. This markedly improves maths, logic, multi-step planning, and debugging.
Reasoning content#
When a model exposes its thinking, it arrives in a separate field so it never contaminates the answer:
{
"choices": [{
"message": {
"role": "assistant",
"content": "The speed is 22.22 m/s.",
"reasoning_content": "120 km in 1.5 hours. 120 km = 120000 m. 1.5 h = 5400 s. 120000 / 5400 = 22.22 m/s."
},
"finish_reason": "stop"
}]
}
content holds the answer to show your user. reasoning_content is optional — display it in a collapsible panel or ignore it.
While streaming, reasoning arrives as delta.reasoning_content and typically completes before the first delta.content chunk.
Token budget#
This is the single most common mistake with reasoning models.
Reasoning tokens count toward completion_tokens. A tight max_tokens can be spent entirely on reasoning, leaving an empty content.
We guard against this: max_tokens below 2048 is raised to 2048, and omitting it defaults to 8192. For hard problems, allow 16000 or more.
| Task | Suggested max_tokens |
|---|---|
| Short answers | 2048 |
| General use | 8192 (default) |
| Complex reasoning, long code | 16000+ |
Choosing a model#
Reasoning models are slower and cost more, since you pay for reasoning tokens too. Prefer them for algorithm design, debugging, maths, and multi-step planning; prefer non-reasoning models for summarising, formatting, translation, and simple lookups.
Example#
from openai import OpenAI
client = OpenAI(api_key="cc_your_key_here", base_url="https://cleanapis.com/v1")
response = client.chat.completions.create(
model="claude-opus-4.8",
messages=[{
"role": "user",
"content": "A train covers 120 km in 1.5 hours. What is its speed in m/s?",
}],
max_tokens=8192,
)
message = response.choices[0].message
print("Answer:", message.content)
reasoning = getattr(message, "reasoning_content", None)
if reasoning:
print("\nThinking:", reasoning)
Still stuck?
Open a support ticket from your dashboard and we'll take a look.