OpenAI-Compatible APIs — Why They Matter and How to Switch
One API shape, many providers. What OpenAI compatibility actually guarantees, where it breaks, and how to migrate an existing integration in two lines.
Contents
- What the standard covers
- Migrating an existing integration
- Why this standardisation is valuable
- Where implementations actually diverge
- 1. Streaming that dies early
- 2. Missing capability metadata
- 3. Usage omitted from streams
- 4. Empty JSON schemas rejected
- 5. Errors that are not JSON
- 6. No CORS
- 7. Anthropic-style clients
- What to test before you commit
- When compatibility is not enough
- Running several providers at once
- Next steps
"OpenAI-compatible" has become the de facto standard for AI model APIs. Dozens of providers implement it, and every serious tool speaks it. That standardisation is quietly one of the most useful things to happen to AI development.
But compatibility is a spectrum, not a binary. This explains what it actually guarantees, where implementations diverge, and how to migrate without rewriting anything.
What the standard covers#
OpenAI compatibility means a provider implements the same HTTP surface as OpenAI's API:
| Endpoint | Purpose |
|---|---|
POST /v1/chat/completions |
Generate a completion |
GET /v1/models |
List available models |
GET /v1/models/{id} |
Retrieve one model |
POST /v1/embeddings |
Create embeddings |
Same request bodies, same response shapes, same streaming format, same error envelope. Which means the official OpenAI SDKs work unchanged.
Migrating an existing integration#
Two lines. That is the entire point.
Before:
from openai import OpenAI
client = OpenAI(api_key="sk-...")
After:
from openai import OpenAI
client = OpenAI(
api_key="cc_your_key_here",
base_url="https://cleanapis.com/v1",
)
Everything downstream is untouched: your prompts, your streaming handler, your retry logic, your token accounting, your error handling.
Same in Node:
const client = new OpenAI({
apiKey: "cc_your_key_here",
baseURL: "https://cleanapis.com/v1",
});
Go, .NET, Java, Ruby — every official SDK exposes a base URL override, because OpenAI itself needs it for Azure deployments.
Why this standardisation is valuable#
No vendor lock-in at the code level. Switching providers is a config change, not a migration project. That changes your negotiating position and your risk profile.
Every tool works everywhere. Kilo Code, Cline, Cursor, opencode, Continue, Aider — none of them integrate with providers individually. They integrate with the shape, and any compliant provider works.
Model choice becomes a string. Trying a different model does not mean a different SDK, different auth, or different response parsing.
Existing knowledge transfers. Documentation, Stack Overflow answers, and tutorials written for OpenAI apply directly.
Where implementations actually diverge#
This is the part marketing pages skip. "OpenAI-compatible" providers differ in ways that break real integrations.
1. Streaming that dies early#
Many providers technically accept "stream": true but close the connection after a few minutes. Fine for chat, fatal for agents mid-refactor.
Ask: how long does a stream stay open, and are keep-alives sent?
Clean APIs: up to one hour, with a : ping comment every 15 seconds.
2. Missing capability metadata#
GET /v1/models returns model IDs everywhere. Whether it returns capabilities varies — and tools depend on it.
Kilo Code reads architecture.input_modalities before attaching a screenshot. Without that field it assumes text-only and refuses, even on a model that handles images perfectly.
We return the full shape:
{
"id": "claude-opus-4.8",
"context_length": 128000,
"architecture": {
"modality": "text+image->text",
"input_modalities": ["text", "image"],
"output_modalities": ["text"]
},
"supported_parameters": ["tools", "tool_choice", "reasoning", "max_tokens"],
"capabilities": ["reasoning", "vision", "tools", "streaming"]
}
3. Usage omitted from streams#
Non-streaming responses include usage. Streaming ones only do if the provider asks the upstream for it. Without it you are estimating token counts, which drifts badly on reasoning models where most output is invisible.
We request it, so the final chunk before [DONE] carries real counts.
4. Empty JSON schemas rejected#
A tool with no parameters is described as:
"parameters": {"type": "object", "properties": {}, "required": []}
Some providers reject that with Invalid JSON schema: [] is not of type "object" — a serialisation bug where {} becomes []. Kilo Code sends properties on every tool, so it triggers this immediately.
5. Errors that are not JSON#
Send a malformed request without an Accept: application/json header and some providers return an HTML error page. Clients that assume JSON crash on it.
Our /v1/* surface always returns the OpenAI error envelope, regardless of headers:
{
"error": {
"message": "The model 'nope' does not exist or you do not have access to it.",
"type": "invalid_request_error",
"code": "model_not_found",
"param": "model"
}
}
6. No CORS#
Browser-based tools like OpenClaw make requests from the page. Without CORS headers on the API, the browser blocks them and you need a proxy.
We send them on /v1/*. Authentication is Bearer-token based rather than cookie based, so a permissive origin policy carries no session risk.
7. Anthropic-style clients#
Claude Code sends x-api-key instead of Authorization: Bearer. Providers that only read the Bearer header return 401, so people build translation proxies.
We accept both.
→ Claude Code with a custom API
What to test before you commit#
Do not trust a compatibility claim. Verify it in ten minutes.
1. Models with capabilities
curl https://cleanapis.com/v1/models \
-H "Authorization: Bearer cc_your_key_here"
Look for capabilities, context_length, and architecture.
2. A basic completion
curl https://cleanapis.com/v1/chat/completions \
-H "Authorization: Bearer cc_your_key_here" \
-H "Content-Type: application/json" \
-d '{"model":"claude-opus-4.8","messages":[{"role":"user","content":"hi"}]}'
3. Streaming with usage
Add "stream": true and -N. Confirm chunks arrive incrementally, usage appears in the final chunk, and the stream ends with data: [DONE].
4. Tool calling, including an empty schema
curl https://cleanapis.com/v1/chat/completions \
-H "Authorization: Bearer cc_your_key_here" \
-H "Content-Type: application/json" \
-d '{
"model":"claude-opus-4.8",
"messages":[{"role":"user","content":"list the files"}],
"tools":[{"type":"function","function":{"name":"list_files","description":"List files","parameters":{"type":"object","properties":{},"required":[]}}}]
}'
If that returns tool_calls, the empty-schema handling is correct.
5. Error shape without an Accept header
Send a request missing messages and confirm you get JSON, not HTML.
When compatibility is not enough#
A few things are genuinely provider-specific and no amount of compatibility fixes them:
Model availability. Compatibility does not mean the same models. Check the catalogue.
Rate limits. Structure and generosity vary. Ours are per-minute on requests and tokens, reported in X-RateLimit-* headers on every response.
Pricing. Same model, different margin.
Payment methods. Most assume a US credit card. We support Stripe, PayPal, bKash, Nagad, and crypto.
Running several providers at once#
Because the shape is standard, a fallback is trivial:
PROVIDERS = [
{"base_url": "https://cleanapis.com/v1", "key": PRIMARY_KEY},
{"base_url": "https://backup.example/v1", "key": BACKUP_KEY},
]
def complete(messages, model):
last_error = None
for p in PROVIDERS:
try:
client = OpenAI(api_key=p["key"], base_url=p["base_url"])
return client.chat.completions.create(model=model, messages=messages)
except Exception as e:
last_error = e
continue
raise last_error
No adapters, no abstraction layer. That is the value of a standard.
Next steps#
Get a free API key — 5M tokens monthly, 31 models, two lines to switch.
Ready to build?
Everything in this article works on the free tier — 5M tokens every month, all 33 models, no card.