Introducing Clean APIs — One Endpoint for Every AI Model
One OpenAI-compatible API for 31 models, from 5M free tokens a month to unlimited. Streaming, tool calling, vision, and reasoning included on every plan.
Contents
- Two lines to switch
- What you actually get
- Every capability, on every plan
- Real streaming, built for agents
- Capability discovery that tools understand
- Pricing that rewards volume
- Payments that work where you are
- Works with the tools you already use
- Transparency, by default
- Why we built it this way
- What we are not
- Start now
Every developer building with AI hits the same wall. You start with one provider, wire up their SDK, and ship. Then a better model launches somewhere else — different SDK, different auth, different response shape, different billing dashboard. Six months later you are maintaining four integrations and reconciling four invoices.
Clean APIs exists to end that. One endpoint, one API key, one bill, and access to 31 models including flagship reasoning models. If your code already talks to OpenAI, it already talks to us.
Two lines to switch#
There is no Clean APIs SDK to learn, because you do not need one. Point any official OpenAI client at our base URL:
from openai import OpenAI
client = OpenAI(
api_key="cc_your_key_here",
base_url="https://cleanapis.com/v1", # ← the only change
)
response = client.chat.completions.create(
model="claude-opus-4.8",
messages=[{"role": "user", "content": "Hello!"}],
)
print(response.choices[0].message.content)
Same in Node:
import OpenAI from "openai";
const client = new OpenAI({
apiKey: "cc_your_key_here",
baseURL: "https://cleanapis.com/v1",
});
const response = await client.chat.completions.create({
model: "claude-opus-4.8",
messages: [{ role: "user", content: "Hello!" }],
});
Request shapes, response shapes, streaming chunks, tool call payloads, and error envelopes all match the OpenAI specification. Your retry logic, your error handling, your token accounting — none of it changes.
What you actually get#
Every capability, on every plan#
Plans differ in token volume only. They never gate features or models. The free tier reaches the same 31 models as the top tier:
| Capability | What it means |
|---|---|
| Streaming | Server-sent events, keep-alive pings, connections held up to an hour |
| Tool calling | Full round trips, parallel calls, streamed argument deltas |
| Vision | Base64 or URL image input on vision-capable models |
| Reasoning | Chain-of-thought returned in a separate field, never mixed into your answer |
| JSON mode | Structured output via response_format |
Real streaming, built for agents#
Most "OpenAI compatible" proxies technically support stream: true but drop the connection after a few minutes. That breaks long agentic coding sessions in the middle of a refactor.
Our streaming connections stay open for up to one hour, with a keep-alive comment every 15 seconds so intermediaries do not time out an idle channel. If the upstream provider fails mid-stream, you get a clean error frame followed by [DONE] — not a truncated response your parser chokes on.
Capability discovery that tools understand#
GET /v1/models returns capability metadata in the format agent tooling expects:
{
"id": "claude-opus-4.8",
"context_length": 128000,
"architecture": {
"modality": "text+image->text",
"input_modalities": ["text", "image"],
"output_modalities": ["text"]
},
"supported_parameters": ["tools", "tool_choice", "reasoning", "max_tokens", "temperature"],
"capabilities": ["reasoning", "vision", "tools", "streaming"]
}
This matters more than it sounds. Tools like Kilo Code read architecture.input_modalities before deciding whether they may attach a screenshot. Without it, they refuse to send images even to a model that handles them perfectly.
Pricing that rewards volume#
Start free. Upgrade when the numbers justify it.
| Plan | Tokens / month | Price | Per 1M tokens |
|---|---|---|---|
| Free | 5M | $0 | — |
| Starter | 50M | $3.50 | $0.070 |
| Basic | 100M | $5.75 | $0.0575 |
| Pro | 200M | $9.50 | $0.0475 |
| Scale | 500M | $15.50 | $0.031 |
| Unlimited | No cap | $50 | — |
Tokens cover prompt and completion combined. Your monthly allowance is spent first; anything beyond it bills against your account balance at each model's published per-token rate.
Run out with an empty balance and you get a clear 402 before the request reaches a model. No surprise invoices, no silent truncation mid-response.
Payments that work where you are#
Most AI APIs assume you hold a US credit card. We support Stripe, PayPal, bKash, Nagad, and crypto (BTC/USDT).
For bKash and Nagad: send the payment, submit the transaction ID, and your plan activates once confirmed. No card, no currency conversion headache.
Works with the tools you already use#
Any client that accepts a custom OpenAI-compatible endpoint works. These are verified end to end:
- Kilo Code — VS Code agent
- Cline — VS Code agent
- Cursor — AI-first editor
- opencode — terminal agent
- Claude Code — Anthropic's CLI agent
- OpenClaw — browser-based agent
We accept the API key as either Authorization: Bearer or x-api-key, so Anthropic-style clients connect without a translation layer.
Transparency, by default#
Every request is logged with its exact token count, latency, cost, and which budget paid for it. Not estimates — the real numbers the provider reported.
You can see, per API key and per model:
- Request count and error rate
- Prompt, completion, and total tokens
- Exact cost in USD, to six decimal places
- Whether each request came from your plan allowance or your balance
Scoped API keys mean a key can be limited to models:read, inference, or embeddings — useful when a browser-side tool needs a key you would rather not give full access.
Why we built it this way#
Three decisions shaped the product, and they are worth stating plainly because they are the reasons to choose us over a thinner proxy.
Compatibility over invention. We could have designed a nicer API. Nobody wants to learn one. Matching the OpenAI shape exactly means every SDK, every tool, every Stack Overflow answer, and every tutorial already applies.
Capability metadata is not optional. A model list without capabilities forces every tool to guess. Guessing means refusing to send images to vision models, or sending tools to models that cannot use them. We return architecture, supported_parameters, and capabilities on every model because that is what makes agent tooling work rather than half-work.
Never gate features by tier. Plans differ in token volume, full stop. Reasoning, vision, tool calling, and streaming are on the free tier. Gating capabilities forces people to upgrade to evaluate, which is a bad way to earn a customer.
What we are not#
Worth being direct about the boundaries.
We are not a model trainer. We do not build or fine-tune models. We route to providers and handle the plumbing — auth, billing, capability discovery, streaming stability, error normalisation.
We are not a hosting platform. There is no deploy step, no runtime, no containers. You call an endpoint.
We are not free forever at scale. 5M tokens a month genuinely is free with no card. Past that, you pay — because upstream inference costs real money and a provider that pretends otherwise eventually disappears.
Start now#
curl https://cleanapis.com/v1/chat/completions \
-H "Authorization: Bearer cc_your_key_here" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-opus-4.8",
"messages": [{"role": "user", "content": "Hello!"}]
}'
Create a free account — 5M tokens every month, no card required. Then read the quickstart or jump straight to connecting your coding agent.
Ready to build?
Everything in this article works on the free tier — 5M tokens every month, all 33 models, no card.