Engineering

Tool Calling Explained — How AI Agents Actually Edit Your Code

The mechanism behind every AI coding agent, explained with working code. Full round trips, parallel calls, streaming deltas, and the empty-schema bug that breaks agents.

6 min read Clean APIs Team
Tool Calling Explained — How AI Agents Actually Edit Your Code
Contents

Every AI coding agent — Kilo Code, Cline, Cursor Composer, opencode, Claude Code — runs on the same mechanism: tool calling. The model does not edit your files. It asks your program to, and your program does it.

Understanding this makes agent behaviour predictable instead of magical, and makes debugging them tractable.

The core idea#

A language model produces text. It cannot read your disk, run a command, or call an API. Tool calling bridges that gap with a structured conversation:

  1. You describe functions the model may request
  2. The model replies "call read_file with path: src/app.py"
  3. Your code executes it and sends the result back
  4. The model continues with that information

The model never touches anything. It only ever asks.

A complete round trip#

Step 1: Describe your tools#

{
  "model": "claude-opus-4.8",
  "messages": [{"role": "user", "content": "What is the weather in Dhaka?"}],
  "tools": [{
    "type": "function",
    "function": {
      "name": "get_weather",
      "description": "Get the current weather for a city",
      "parameters": {
        "type": "object",
        "properties": {
          "city": {"type": "string", "description": "City name"},
          "unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
        },
        "required": ["city"]
      }
    }
  }],
  "tool_choice": "auto"
}

The description fields matter more than people expect — they are how the model decides whether and how to call the function. Vague descriptions produce wrong calls.

Step 2: The model requests a call#

{
  "choices": [{
    "message": {
      "role": "assistant",
      "content": null,
      "tool_calls": [{
        "id": "call_abc123",
        "type": "function",
        "function": {
          "name": "get_weather",
          "arguments": "{\"city\":\"Dhaka\",\"unit\":\"celsius\"}"
        }
      }]
    },
    "finish_reason": "tool_calls"
  }]
}

Two details that trip people up:

  • arguments is a JSON string, not an object. Parse it before use.
  • finish_reason is tool_calls, not stop. That is your signal the model is waiting on you.

Step 3: Execute and return the result#

Append the assistant message verbatim, then a tool message carrying the output:

{
  "model": "claude-opus-4.8",
  "messages": [
    {"role": "user", "content": "What is the weather in Dhaka?"},
    {
      "role": "assistant",
      "content": null,
      "tool_calls": [{
        "id": "call_abc123",
        "type": "function",
        "function": {"name": "get_weather", "arguments": "{\"city\":\"Dhaka\"}"}
      }]
    },
    {
      "role": "tool",
      "tool_call_id": "call_abc123",
      "content": "{\"temp_c\":31,\"sky\":\"humid\"}"
    }
  ],
  "tools": [ ... same definitions ... ]
}

The tool_call_id must match. That is how the model pairs your result with its request when several are in flight.

Step 4: The model answers#

It's 31°C and humid in Dhaka right now.

Working Python implementation#

import json
from openai import OpenAI

client = OpenAI(api_key="cc_your_key_here", base_url="https://cleanapis.com/v1")

tools = [{
    "type": "function",
    "function": {
        "name": "get_weather",
        "description": "Get the current weather for a city",
        "parameters": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    },
}]

def get_weather(city: str) -> dict:
    # Your real implementation goes here
    return {"temp_c": 31, "sky": "humid", "city": city}

messages = [{"role": "user", "content": "What is the weather in Dhaka?"}]

response = client.chat.completions.create(
    model="claude-opus-4.8", messages=messages, tools=tools
)
message = response.choices[0].message

if message.tool_calls:
    # Keep the assistant message exactly as returned
    messages.append(message)

    for call in message.tool_calls:
        args = json.loads(call.function.arguments)
        result = get_weather(**args)

        messages.append({
            "role": "tool",
            "tool_call_id": call.id,
            "content": json.dumps(result),
        })

    final = client.chat.completions.create(
        model="claude-opus-4.8", messages=messages, tools=tools
    )
    print(final.choices[0].message.content)
else:
    print(message.content)

How agents use this#

A coding agent defines tools like read_file, write_file, list_directory, run_command, and search_codebase. Then it loops:

User: "Add input validation to the signup endpoint"

→ model calls search_codebase("signup")
→ model calls read_file("routes/auth.py")
→ model reasons about the code
→ model calls write_file("routes/auth.py", <new content>)
→ model calls run_command("pytest tests/test_auth.py")
→ tests fail
→ model reads the failure, calls write_file again
→ tests pass
→ model answers: "Added validation, tests pass"

Every arrow is a full API round trip. That is why agents cost far more than chat — a single instruction can be fifteen requests, each resending the growing conversation.

tool_choice controls the decision#

Value Behaviour
"auto" Model decides (default when tools present)
"none" Never call a tool
"required" Must call at least one
{"type": "function", "function": {"name": "get_weather"}} Must call that one

"required" is useful when you know a tool is needed and want to skip a turn of the model deciding.

Parallel calls#

A single response can contain several entries in tool_calls. Execute them all and return one tool message per call:

messages.append(message)

for call in message.tool_calls:            # may be more than one
    result = dispatch(call.function.name, json.loads(call.function.arguments))
    messages.append({
        "role": "tool",
        "tool_call_id": call.id,           # each result matched by id
        "content": json.dumps(result),
    })

Agents use this to read three files in one turn instead of three, which meaningfully reduces both latency and token cost.

Streaming tool calls#

With stream: true, tool calls arrive as deltas — and arguments comes in fragments:

data: {"choices":[{"delta":{"tool_calls":[{"index":0,"id":"call_abc","function":{"name":"get_weather","arguments":""}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"{\"ci"}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"ty\":\"Dhaka\"}"}}]}}]}
data: {"choices":[{"delta":{},"finish_reason":"tool_calls"}]}

Accumulate by index and only parse once the stream finishes:

calls = {}

for chunk in stream:
    for delta in (chunk.choices[0].delta.tool_calls or []):
        c = calls.setdefault(delta.index, {"id": None, "name": None, "args": ""})
        if delta.id:
            c["id"] = delta.id
        if delta.function.name:
            c["name"] = delta.function.name
        if delta.function.arguments:
            c["args"] += delta.function.arguments      # fragments!

# Now the JSON is complete
for c in calls.values():
    args = json.loads(c["args"])

Parsing a partial arguments string is the most common streaming bug. It is only valid JSON at the end.

The empty-schema bug#

Here is a real failure worth knowing, because it breaks agents in a way that looks like a provider problem.

Tools with no parameters are described like this:

"parameters": { "type": "object", "properties": {}, "required": [] }

Note properties is an empty object — {}.

In PHP, json_decode($body, true) turns both {} and [] into the same empty array. Re-encoding then emits [], and the provider rejects it:

Invalid JSON schema: [] is not of type "object"

Kilo Code sends properties on every tool, including parameterless ones like list_files, so it hits this immediately while other clients that omit the key do not.

Clean APIs reads tools, response_format, and messages from the raw request body rather than the array-decoded version, preserving the object/array distinction. Empty schemas pass through intact. It is the kind of bug that only shows up with real agent traffic.

Which models support it#

Tool calling is a capability, not universal. Check the 🔧 Tools badge on the models page, or:

curl https://cleanapis.com/v1/models \
  -H "Authorization: Bearer cc_your_key_here"
{
  "capabilities": ["reasoning", "vision", "tools", "streaming"],
  "supported_parameters": ["tools", "tool_choice", "max_tokens"]
}

Without "tools", an agent pointed at that model can talk but cannot act. Every Clean APIs plan reaches every tool-capable model — the capability is never gated by tier.

Cost implications#

Two things inflate agent token usage beyond what people expect.

Tool schemas ride along on every request. Five verbose definitions can be 1,500 tokens, resent on every turn of the conversation. Twenty turns is 30,000 tokens spent purely repeating definitions.

Keep descriptions tight, and send only the tools relevant to the current task.

Every step is a full round trip with the whole history. A fifteen-step task does not cost 15× a single request — it costs more, because the history grows each time.

→ Understanding token pricing

Next steps#

Get a free API key — 5M tokens monthly, tool calling on every plan.

Ready to build?

Everything in this article works on the free tier — 5M tokens every month, all 33 models, no card.

Related reading