AI Coding Tools

Using Claude Code with a Custom API Endpoint

Point Claude Code at any Anthropic-compatible API using environment variables. No proxy required, plus model selection and troubleshooting.

6 min read Clean APIs Team
Using Claude Code with a Custom API Endpoint
Contents

Claude Code is Anthropic's official CLI agent, and it is very good at long autonomous tasks. It speaks the Anthropic protocol rather than OpenAI's, which is why most "OpenAI-compatible" providers require a translation proxy to work with it.

Clean APIs does not. We accept the API key as x-api-key — the header Anthropic-style clients send — alongside Authorization: Bearer. Point Claude Code at us with three environment variables and it works.

What you need#

  • Claude Code installed
  • A Clean APIs API key — free, 5M tokens/month, no card
  • Two minutes

Step 1: Get your API key#

Sign up at cleanapis.com, then Dashboard → API Keys → Create Key. Keep all scopes enabled.

Step 2: Set three environment variables#

export ANTHROPIC_BASE_URL=https://cleanapis.com/v1
export ANTHROPIC_AUTH_TOKEN=cc_your_key_here
export ANTHROPIC_MODEL=claude-opus-4.8

claude

On Windows PowerShell:

$env:ANTHROPIC_BASE_URL = "https://cleanapis.com/v1"
$env:ANTHROPIC_AUTH_TOKEN = "cc_your_key_here"
$env:ANTHROPIC_MODEL = "claude-opus-4.8"

claude

Add the exports to your ~/.bashrc, ~/.zshrc, or PowerShell profile to make them permanent.

Step 3: Verify#

Start Claude Code in a project and ask something small:

What does this project do? Read the README and summarise it.

If it reads the file and answers, the connection works.

Why no proxy is needed#

Anthropic's API and OpenAI's differ in two places that matter for a client like this:

Authentication header. Anthropic clients send x-api-key: sk-…; OpenAI clients send Authorization: Bearer sk-…. We accept either, so no header rewriting is required.

Request shape. Claude Code sends OpenAI-compatible chat payloads when pointed at a custom base URL, which our /v1/chat/completions handles natively.

That is the whole reason this works with plain environment variables while other providers need a shim.

Choosing a model#

Claude Code is an autonomous agent, so the same three criteria apply as for any agent:

Tool calling is mandatory. Without it, Claude Code cannot read or edit files. Check the Tools badge on the models page.

Large context. It sends file contents plus its own history. Prefer 128K or more.

Reasoning for hard problems. Claude Code works through multi-step tasks; a reasoning model measurably helps on debugging and architecture. Costs more, because reasoning output is billed.

Set ANTHROPIC_MODEL to any ID from:

curl https://cleanapis.com/v1/models \
  -H "x-api-key: cc_your_key_here"

Note that works with x-api-key too — useful for verifying the Anthropic-style path directly.

Switching models per session#

Because it is an environment variable, changing models is a shell command:

# Cheap model for mechanical work
ANTHROPIC_MODEL=some-fast-model claude

# Reasoning model for a hard bug
ANTHROPIC_MODEL=claude-opus-4.8 claude

Every Clean APIs plan reaches every model, so switching has no billing implications beyond the per-token rate.

A useful pattern is a shell alias per role:

alias claude-fast='ANTHROPIC_MODEL=some-fast-model claude'
alias claude-deep='ANTHROPIC_MODEL=claude-opus-4.8 claude'

Long autonomous sessions#

Claude Code is designed to work unattended for a long time, which makes connection stability matter more than with interactive tools.

Our side: streaming connections held open up to an hour, keep-alive comment every 15 seconds so intermediaries do not drop an idle channel.

Your side: if you route through your own proxy, match the timeout:

proxy_read_timeout 3700s;
proxy_buffering off;

Run it inside tmux or screen on remote machines so a dropped SSH session does not kill an in-progress task.

Troubleshooting#

401 Unauthorized#

The environment variable is not set in the shell Claude Code is running in:

echo $ANTHROPIC_AUTH_TOKEN

Empty means the export did not persist — add it to your shell profile. Note the variable is ANTHROPIC_AUTH_TOKEN, not ANTHROPIC_API_KEY.

404 model_not_found#

ANTHROPIC_MODEL does not match a real ID. List them with the curl command above and copy the id exactly.

403 insufficient_scope#

The key lacks a scope. Enable models:read and inference.

Empty response#

A reasoning model spent its entire token budget on internal reasoning. We floor max_tokens at 2048 and default to 8192, but a hard problem may need 16000+.

It still calls Anthropic directly#

ANTHROPIC_BASE_URL is not set, or was set in a different shell. Verify:

echo $ANTHROPIC_BASE_URL

It must print https://cleanapis.com/v1.

429 rate limited#

Per-minute limit hit. Claude Code can burst hard on long tasks. Retry-After gives the wait; Dashboard → Usage → Rate Limits shows live consumption.

402 insufficient balance#

Allowance and balance both empty. Nothing charged for the rejected request. Top up or upgrade.

Cost awareness#

Autonomous agents are the most expensive way to use an AI API, because they loop: read, reason, act, verify, repeat.

Task Tokens
Single file edit ~8,000
Multi-file refactor 30,000–80,000
Long autonomous session 100,000+

Dashboard → Usage logs every request with its real token count and cost to six decimal places. Set a usage alert in Settings → Notifications to be warned at whatever percentage of your allowance you choose.

→ Understanding token pricing

Getting good results from long autonomous runs#

Claude Code is at its best when given a clear objective and left alone. Getting there reliably takes a bit of setup.

Give it a verification command. The single biggest quality improvement is telling it how to check its own work: "implement this, then run pytest tests/ until it passes." Without a verification loop it writes code and stops, and you find the bugs. With one, it finds them itself.

Scope the task narrowly. "Refactor the auth module" invites an unbounded run. "Extract token validation from ApiKeyAuth into a testable service, keeping the existing behaviour, and confirm with the existing tests" gives it a finish line.

Commit before starting. Autonomous agents edit files directly. A clean Git state means git diff is your review and git checkout . is your undo.

Watch the first few minutes. If it is heading the wrong way, an early interrupt costs almost nothing. Twenty minutes in, you have paid for a lot of tokens going nowhere.

Prefer a large context model. Long runs accumulate history — files read, commands run, results considered. A small window forces it to forget what it already learned, and it starts repeating work.

Cost characteristics of autonomous runs#

Autonomous agents have a different cost profile from interactive tools, and it is worth internalising.

Interactive tools cost roughly in proportion to your typing. You ask, it answers, you read. Autonomous agents cost in proportion to task difficulty and verification loops — and neither is visible when you start.

A task that works first try might be 20,000 tokens. The same task where tests fail twice and the model reconsiders its approach can be 150,000. That variance is inherent to the approach, not a defect.

Two practical responses:

Set a low-balance alert. In Settings → Notifications, so you find out before a run stops mid-task rather than after.

Check usage after the first few tasks. Dashboard → Usage shows real per-request costs. After a handful of runs you will know what a typical task costs on your codebase, which makes plan sizing an informed decision instead of a guess.

Next steps#

Get a free API key — 5M tokens monthly, 31 models, no card.

Ready to build?

Everything in this article works on the free tier — 5M tokens every month, all 33 models, no card.

Related reading