AI Coding Tools

How to Use a Custom AI API with Cursor (Step-by-Step)

Override Cursor's OpenAI base URL to use any model provider. Full setup, what works and what does not, and how to keep costs predictable.

6 min read Clean APIs Team
How to Use a Custom AI API with Cursor (Step-by-Step)
Contents

Cursor has the best inline AI editing experience of any tool — tab completion that predicts multi-line changes, and a fast selection-edit flow that feels native because the AI is built into the editor rather than bolted on.

It also lets you override the OpenAI endpoint, which means you are not restricted to the models it ships with. This guide connects Cursor to Clean APIs — 31 models, one key, starting free.

What you need#

  • Cursor installed (cursor.com)
  • A Clean APIs API key — free, 5M tokens/month, no card
  • Three minutes

Step 1: Get your API key#

Sign up at cleanapis.com, then Dashboard → API Keys → Create Key. Keep inference and models:read enabled.

Copy the key — starts with cc_, shown once.

Step 2: Override the base URL in Cursor#

  1. Settings → Models (or Cmd/Ctrl + Shift + J → Models)
  2. Scroll to OpenAI API Key
  3. Paste your cc_… key
  4. Enable Override OpenAI Base URL
  5. Enter: https://cleanapis.com/v1
  6. Click Verify

Verification calls /v1/models, which is why the key needs models:read.

Step 3: Add a model#

Cursor does not auto-discover models — you type the exact ID.

  1. In Settings → Models, click + Add model
  2. Enter a model ID exactly as returned by our API, e.g. claude-opus-4.8
  3. Save, then select it in the chat model picker

Get the exact IDs:

curl https://cleanapis.com/v1/models \
  -H "Authorization: Bearer cc_your_key_here"

Copy the id field verbatim. A typo produces 404 model_not_found.

What works with a custom endpoint#

Chat panel. Full conversations with your chosen model, including streaming.

Selection edits. Highlight code, describe the change, apply. Works normally.

Composer / multi-file. Works with models that support tool calling.

Streaming. Responses stream token by token as expected.

What does not#

Tab completion. Cursor's autocomplete uses its own proprietary model and is not routed through a custom endpoint. That feature stays on Cursor's infrastructure regardless of your API settings.

Codebase indexing. Also Cursor's own system.

This is worth understanding before you switch: overriding the endpoint changes chat and edits, not autocomplete. If tab completion is the main reason you use Cursor, you will still be on their subscription for that.

Choosing a model#

For chat and reasoning: a reasoning-capable model earns its cost on debugging and architecture questions. Set a generous max_tokens — reasoning output is billed and a small cap can consume the whole budget on thinking, returning an empty answer.

For selection edits: a fast mid-tier model. These are short, focused requests where latency matters more than depth.

For Composer: must have the tools capability, or multi-file editing cannot function.

Filter by capability on the models page. Every Clean APIs plan reaches every model, so you can keep several configured and switch per task.

Cost: read this before switching#

Cursor's subscription and a model API are separate costs. Overriding the endpoint does not reduce your Cursor bill — it moves chat and edit usage onto your own API budget.

That is worth it when:

  • You want models Cursor does not offer
  • You are hitting Cursor's usage caps
  • You want a single usage dashboard across all your tools
  • You need models available in your region

It is not worth it if you are happy with the built-in models and only use autocomplete.

Realistic numbers:

Usage Tokens
Chat question about code ~3,000
Selection edit ~2,000
Composer multi-file change 20,000–60,000

Dashboard → Usage shows the real numbers per request — token counts, latency, cost to six decimal places.

→ Understanding token pricing

Troubleshooting#

"Verify" fails#

Three usual causes:

  1. Base URL includes /chat/completions — it must end at /v1
  2. Key lacks models:read
  3. Key is revoked

Test independently:

curl https://cleanapis.com/v1/models \
  -H "Authorization: Bearer cc_your_key_here"

Models returned means your key and URL are fine, and the issue is in Cursor's settings.

404 model_not_found#

The model ID does not match. Copy it exactly from GET /v1/models — IDs are case-sensitive and include punctuation.

Empty or truncated responses#

A reasoning model spent its budget thinking. Increase max_tokens in Cursor's model settings; 16000+ for hard problems.

429 rate limited#

Per-minute limit hit. Retry-After says how long. Composer can burst hard — Dashboard → Usage → Rate Limits shows live consumption.

402 insufficient balance#

Allowance and balance both empty. Nothing charged for the rejection. Top up or upgrade.

Autocomplete still uses Cursor's model#

Expected. See "What does not work" above.

Running Cursor alongside other tools#

A common setup:

  • Cursor for inline editing and its autocomplete
  • Kilo Code or Cline in VS Code for autonomous multi-file work
  • One Clean APIs key powering the model side of both

Since one key covers every tool, there is nothing to reconcile — usage from all of them lands in the same dashboard, split by key so you can see which tool spent what.

Create a separate key per tool if you want that breakdown.

Choosing which work goes where#

Once you have both a custom endpoint and Cursor's built-in models available, the split that works in practice:

Cursor's own models for autocomplete — you cannot route that anyway — and for quick inline edits where latency dominates.

Your custom endpoint for chat, longer reasoning, and Composer runs. These are the requests where model choice actually changes the outcome, and where a reasoning model earns its cost.

That means the model picker becomes a deliberate choice rather than a default. A useful habit is naming your added models by role rather than by model family, so the dropdown reads "deep reasoning" and "fast edits" instead of two similar version strings.

Limitations worth knowing before you switch#

Autocomplete stays on Cursor. Already covered above, but it is the single most common surprise, so it bears repeating: overriding the endpoint changes chat and edits only.

No capability introspection. Cursor does not read model capabilities from /v1/models, so it will let you select a text-only model and then fail when you attach an image. Tools like Kilo Code and Cline check first and guide you; Cursor does not.

Manual model list. Every model you want must be typed in by hand. When you add models to your account, Cursor does not notice until you add them there too.

Composer needs tool calling. If multi-file editing silently does nothing, the selected model almost certainly lacks the tools capability. Check the badge on the models page.

None of these are dealbreakers — they are the cost of Cursor's tighter, more opinionated integration. Knowing them up front saves an afternoon of confusion.

Next steps#

Get a free API key — 5M tokens monthly, 31 models, no card.

Ready to build?

Everything in this article works on the free tier — 5M tokens every month, all 33 models, no card.

Related reading