Chat Completions
Full reference for /chat/completions, including all supported parameters.
On this page
Chat Completions#
POST https://cleanapis.com/v1/chat/completions
Request body#
| Field | Type | Required | Description |
|---|---|---|---|
model |
string | Yes | Model ID, e.g. claude-opus-4.8 |
messages |
array | Yes | Conversation history (see below) |
stream |
bool | No | Stream the response as SSE. Default false |
max_tokens |
int | No | Maximum tokens to generate. Defaults to 8192 |
temperature |
float | No | 0.0–2.0. Default 1.0 |
top_p |
float | No | 0.0–1.0 nucleus sampling |
tools |
array | No | Function definitions the model may call |
tool_choice |
string|object | No | auto, none, required, or a specific function |
response_format |
object | No | {"type": "json_object"} for JSON-only output |
stop |
string|array | No | Sequences that halt generation |
frequency_penalty |
float | No | -2.0–2.0 |
presence_penalty |
float | No | -2.0–2.0 |
seed |
int | No | Best-effort deterministic sampling |
Message roles#
| Role | Purpose |
|---|---|
system |
Instructions that shape the model's behaviour |
user |
Input from the end user |
assistant |
A previous model reply, or a message carrying tool_calls |
tool |
The result of a tool call, paired with tool_call_id |
max_tokens note: several models spend tokens on internal reasoning before emitting an answer. A small budget can be consumed entirely by reasoning, leaving no visible content, so we raise anything below 2048 and default to 8192 when omitted.
Response#
{
"id": "chatcmpl-6a881ffef05ea",
"object": "chat.completion",
"created": 1787305982,
"model": "claude-opus-4.8",
"choices": [{
"index": 0,
"message": { "role": "assistant", "content": "Hello! How can I help?" },
"finish_reason": "stop"
}],
"usage": {
"prompt_tokens": 10,
"completion_tokens": 8,
"total_tokens": 18
}
}
finish_reason values#
| Value | Meaning |
|---|---|
stop |
Completed naturally or hit a stop sequence |
length |
Ran into max_tokens |
tool_calls |
The model wants you to run a tool |
Multi-turn conversation#
Send the full history each request — the API is stateless.
{
"model": "claude-opus-4.8",
"messages": [
{"role": "system", "content": "You are a concise assistant."},
{"role": "user", "content": "What is Laravel?"},
{"role": "assistant", "content": "A PHP web framework."},
{"role": "user", "content": "Who created it?"}
]
}
Billing#
You are charged for prompt_tokens and completion_tokens at the model's per-1K rates, shown on the Models page. Free plan allowance is consumed first, then your USD balance. Streamed requests are billed after the stream completes.
Still stuck?
Open a support ticket from your dashboard and we'll take a look.