AI Coding Tools

How to Use CleanAPIs with Kilo Code (Complete 2026 Setup Guide)

Connect Kilo Code to 31 AI models with one API key. Full setup, model selection for agentic work, vision support, and fixes for every common error.

5 min read Clean APIs Team
How to Use CleanAPIs with Kilo Code (Complete 2026 Setup Guide)
Contents

Kilo Code is one of the strongest open-source AI coding agents for VS Code. It reads your codebase, edits files, runs commands, and works through multi-step tasks on its own. What it does not do is lock you into one model provider.

This guide connects Kilo Code to Clean APIs, giving it access to 31 models through a single API key — including flagship reasoning models — starting free.

What you need#

  • VS Code with the Kilo Code extension installed
  • A Clean APIs API key — create one free, 5M tokens/month, no card
  • Five minutes

Step 1: Get your API key#

Sign up at cleanapis.com, then open Dashboard → API Keys → Create Key.

Keep all three scopes enabled:

Scope Why Kilo Code needs it
models:read It calls GET /v1/models before the first message to discover capabilities
inference Chat completions — the actual work
embeddings Only if you use embedding features

If you disable models:read, Kilo Code cannot enumerate models and shows an empty dropdown. That is the single most common misconfiguration.

Copy the key — it starts with cc_ and is shown once.

Step 2: Point Kilo Code at Clean APIs#

  1. Open the Kilo Code panel in VS Code
  2. Click the settings gear in the panel header
  3. Set these four fields:
Setting Value
API Provider OpenAI Compatible
Base URL https://cleanapis.com/v1
API Key your cc_… key
Model claude-opus-4.8
  1. Click Done, then start a new task

Note the base URL ends at /v1 — not /v1/chat/completions. Kilo Code appends the path itself. This trips up a lot of people.

Step 3: Verify it works#

Type something small into a new Kilo Code task:

List the files in this project and tell me what framework it uses.

If Kilo Code reads your directory and answers, the connection works. That single request exercised authentication, model resolution, streaming, and tool calling all at once.

Choosing the right model for agentic work#

Not every model suits an autonomous agent. Three things matter, in this order:

1. Tool calling is non-negotiable#

Without the tools capability, Kilo Code can talk about your code but cannot edit files or run commands. Check the Tools badge on the models page, or read it programmatically:

curl https://cleanapis.com/v1/models \
  -H "Authorization: Bearer cc_your_key_here"

Look for "tools" in the capabilities array.

2. Context window decides how much code fits#

Agents send whole files, sometimes several at once, plus their own reasoning history. A 8K-context model runs out almost immediately on a real codebase.

Prefer 128K context or larger for anything beyond single-file edits.

3. Reasoning helps on hard problems, costs more#

Reasoning models think before answering, which measurably improves debugging, architecture decisions, and algorithm work. They also spend tokens on that thinking, so they cost more per task.

A practical split:

Task Model choice
Renaming, formatting, boilerplate Cheap non-reasoning model
Feature implementation, refactoring Mid-tier with tools
Debugging, architecture, algorithms Reasoning model, high max_tokens

Because every Clean APIs plan reaches every model, you can switch per task without touching billing.

Sending screenshots to Kilo Code#

Kilo Code can accept images — UI mockups, error screenshots, diagrams — but only when the selected model advertises vision support.

It checks architecture.input_modalities from GET /v1/models before attaching an image. We return that field, so vision-capable models are detected automatically:

"architecture": {
  "modality": "text+image->text",
  "input_modalities": ["text", "image"],
  "output_modalities": ["text"]
}

If images are refused, the model you picked has no vision capability. Switch models — do not change settings.

Long sessions without timeouts#

Big refactors run for a long time. Two things break them on most providers:

Provider timeouts. Our streaming connections stay open up to one hour, with a keep-alive comment every 15 seconds so intermediaries do not drop an idle channel.

Cancelled requests losing accounting. When you interrupt Kilo Code mid-stream, tokens the provider already generated are still billed and logged. That is deliberate: you see exactly what you spent, and we do not absorb upstream cost silently.

If you self-host behind a reverse proxy, raise its timeout to match:

# Apache
Timeout 3700
ProxyTimeout 3700
# nginx
proxy_read_timeout 3700s;
proxy_buffering off;

proxy_buffering off matters on nginx — with buffering on, streamed chunks are held back and Kilo Code appears frozen.

Troubleshooting#

"Response ended unexpectedly"#

Something between you and us is buffering the stream. Check that X-Accel-Buffering: no is not being stripped, and that nginx has proxy_buffering off.

Empty response from a reasoning model#

Reasoning tokens count toward the completion budget, so a small max_tokens can be consumed entirely by thinking. We raise anything under 2048 and default to 8192, but a hard problem may need 16000+. Set it higher in Kilo Code's advanced settings.

401 Unauthorized#

The key is missing, malformed, or revoked. Keys start with cc_ and are shown once at creation — if you lost it, use Rotate in Dashboard → API Keys to get a replacement.

403 insufficient_scope#

The key lacks a scope Kilo Code needs. Enable models:read.

404 model_not_found#

The model ID is wrong or unavailable to you. List valid IDs with GET /v1/models and copy the exact id string.

402 — out of balance#

Plan allowance and balance are both exhausted. Nothing is charged for the rejected request. Top up or upgrade.

429 — rate limited#

You hit your plan's per-minute limit. The Retry-After header says how long to wait; Dashboard → Usage → Rate Limits shows live consumption.

Models dropdown is empty#

Almost always a missing models:read scope, or a base URL that includes /chat/completions. Verify from a terminal:

curl https://cleanapis.com/v1/models \
  -H "Authorization: Bearer cc_your_key_here"

A JSON list means your URL and key are correct, and the problem is in the extension settings.

Watching what it costs#

Agents are token-hungry. A single "refactor this module" task can run 50K+ tokens across many turns.

Dashboard → Usage shows every request with its real token count, latency, and cost to six decimal places — plus whether it came from your plan allowance or your balance.

Set a usage alert in Settings → Notifications to get warned at whatever percentage of your allowance you choose, before you hit the wall.

Next steps#

Create a free API key and try it — 5M tokens monthly, all 31 models, no card.

Ready to build?

Everything in this article works on the free tier — 5M tokens every month, all 33 models, no card.

Related reading