Tool Calling Explained — How AI Agents Actually Edit Your Code
The mechanism behind every AI coding agent, explained with working code. Full round trips, parallel calls, streaming deltas, and the empty-schema bug that breaks agents.
Contents
- The core idea
- A complete round trip
- Step 1: Describe your tools
- Step 2: The model requests a call
- Step 3: Execute and return the result
- Step 4: The model answers
- Working Python implementation
- How agents use this
- tool_choice controls the decision
- Parallel calls
- Streaming tool calls
- The empty-schema bug
- Which models support it
- Cost implications
- Next steps
Every AI coding agent — Kilo Code, Cline, Cursor Composer, opencode, Claude Code — runs on the same mechanism: tool calling. The model does not edit your files. It asks your program to, and your program does it.
Understanding this makes agent behaviour predictable instead of magical, and makes debugging them tractable.
The core idea#
A language model produces text. It cannot read your disk, run a command, or call an API. Tool calling bridges that gap with a structured conversation:
- You describe functions the model may request
- The model replies "call
read_filewithpath: src/app.py" - Your code executes it and sends the result back
- The model continues with that information
The model never touches anything. It only ever asks.
A complete round trip#
Step 1: Describe your tools#
{
"model": "claude-opus-4.8",
"messages": [{"role": "user", "content": "What is the weather in Dhaka?"}],
"tools": [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name"},
"unit": {"type": "string", "enum": ["celsius", "fahrenheit"]}
},
"required": ["city"]
}
}
}],
"tool_choice": "auto"
}
The description fields matter more than people expect — they are how the model decides whether and how to call the function. Vague descriptions produce wrong calls.
Step 2: The model requests a call#
{
"choices": [{
"message": {
"role": "assistant",
"content": null,
"tool_calls": [{
"id": "call_abc123",
"type": "function",
"function": {
"name": "get_weather",
"arguments": "{\"city\":\"Dhaka\",\"unit\":\"celsius\"}"
}
}]
},
"finish_reason": "tool_calls"
}]
}
Two details that trip people up:
argumentsis a JSON string, not an object. Parse it before use.finish_reasonistool_calls, notstop. That is your signal the model is waiting on you.
Step 3: Execute and return the result#
Append the assistant message verbatim, then a tool message carrying the output:
{
"model": "claude-opus-4.8",
"messages": [
{"role": "user", "content": "What is the weather in Dhaka?"},
{
"role": "assistant",
"content": null,
"tool_calls": [{
"id": "call_abc123",
"type": "function",
"function": {"name": "get_weather", "arguments": "{\"city\":\"Dhaka\"}"}
}]
},
{
"role": "tool",
"tool_call_id": "call_abc123",
"content": "{\"temp_c\":31,\"sky\":\"humid\"}"
}
],
"tools": [ ... same definitions ... ]
}
The tool_call_id must match. That is how the model pairs your result with its request when several are in flight.
Step 4: The model answers#
It's 31°C and humid in Dhaka right now.
Working Python implementation#
import json
from openai import OpenAI
client = OpenAI(api_key="cc_your_key_here", base_url="https://cleanapis.com/v1")
tools = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
def get_weather(city: str) -> dict:
# Your real implementation goes here
return {"temp_c": 31, "sky": "humid", "city": city}
messages = [{"role": "user", "content": "What is the weather in Dhaka?"}]
response = client.chat.completions.create(
model="claude-opus-4.8", messages=messages, tools=tools
)
message = response.choices[0].message
if message.tool_calls:
# Keep the assistant message exactly as returned
messages.append(message)
for call in message.tool_calls:
args = json.loads(call.function.arguments)
result = get_weather(**args)
messages.append({
"role": "tool",
"tool_call_id": call.id,
"content": json.dumps(result),
})
final = client.chat.completions.create(
model="claude-opus-4.8", messages=messages, tools=tools
)
print(final.choices[0].message.content)
else:
print(message.content)
How agents use this#
A coding agent defines tools like read_file, write_file, list_directory, run_command, and search_codebase. Then it loops:
User: "Add input validation to the signup endpoint"
→ model calls search_codebase("signup")
→ model calls read_file("routes/auth.py")
→ model reasons about the code
→ model calls write_file("routes/auth.py", <new content>)
→ model calls run_command("pytest tests/test_auth.py")
→ tests fail
→ model reads the failure, calls write_file again
→ tests pass
→ model answers: "Added validation, tests pass"
Every arrow is a full API round trip. That is why agents cost far more than chat — a single instruction can be fifteen requests, each resending the growing conversation.
tool_choice controls the decision#
| Value | Behaviour |
|---|---|
"auto" |
Model decides (default when tools present) |
"none" |
Never call a tool |
"required" |
Must call at least one |
{"type": "function", "function": {"name": "get_weather"}} |
Must call that one |
"required" is useful when you know a tool is needed and want to skip a turn of the model deciding.
Parallel calls#
A single response can contain several entries in tool_calls. Execute them all and return one tool message per call:
messages.append(message)
for call in message.tool_calls: # may be more than one
result = dispatch(call.function.name, json.loads(call.function.arguments))
messages.append({
"role": "tool",
"tool_call_id": call.id, # each result matched by id
"content": json.dumps(result),
})
Agents use this to read three files in one turn instead of three, which meaningfully reduces both latency and token cost.
Streaming tool calls#
With stream: true, tool calls arrive as deltas — and arguments comes in fragments:
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"id":"call_abc","function":{"name":"get_weather","arguments":""}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"{\"ci"}}]}}]}
data: {"choices":[{"delta":{"tool_calls":[{"index":0,"function":{"arguments":"ty\":\"Dhaka\"}"}}]}}]}
data: {"choices":[{"delta":{},"finish_reason":"tool_calls"}]}
Accumulate by index and only parse once the stream finishes:
calls = {}
for chunk in stream:
for delta in (chunk.choices[0].delta.tool_calls or []):
c = calls.setdefault(delta.index, {"id": None, "name": None, "args": ""})
if delta.id:
c["id"] = delta.id
if delta.function.name:
c["name"] = delta.function.name
if delta.function.arguments:
c["args"] += delta.function.arguments # fragments!
# Now the JSON is complete
for c in calls.values():
args = json.loads(c["args"])
Parsing a partial arguments string is the most common streaming bug. It is only valid JSON at the end.
The empty-schema bug#
Here is a real failure worth knowing, because it breaks agents in a way that looks like a provider problem.
Tools with no parameters are described like this:
"parameters": { "type": "object", "properties": {}, "required": [] }
Note properties is an empty object — {}.
In PHP, json_decode($body, true) turns both {} and [] into the same empty array. Re-encoding then emits [], and the provider rejects it:
Invalid JSON schema: [] is not of type "object"
Kilo Code sends properties on every tool, including parameterless ones like list_files, so it hits this immediately while other clients that omit the key do not.
Clean APIs reads tools, response_format, and messages from the raw request body rather than the array-decoded version, preserving the object/array distinction. Empty schemas pass through intact. It is the kind of bug that only shows up with real agent traffic.
Which models support it#
Tool calling is a capability, not universal. Check the 🔧 Tools badge on the models page, or:
curl https://cleanapis.com/v1/models \
-H "Authorization: Bearer cc_your_key_here"
{
"capabilities": ["reasoning", "vision", "tools", "streaming"],
"supported_parameters": ["tools", "tool_choice", "max_tokens"]
}
Without "tools", an agent pointed at that model can talk but cannot act. Every Clean APIs plan reaches every tool-capable model — the capability is never gated by tier.
Cost implications#
Two things inflate agent token usage beyond what people expect.
Tool schemas ride along on every request. Five verbose definitions can be 1,500 tokens, resent on every turn of the conversation. Twenty turns is 30,000 tokens spent purely repeating definitions.
Keep descriptions tight, and send only the tools relevant to the current task.
Every step is a full round trip with the whole history. A fifteen-step task does not cost 15× a single request — it costs more, because the history grows each time.
Next steps#
Get a free API key — 5M tokens monthly, tool calling on every plan.
Ready to build?
Everything in this article works on the free tier — 5M tokens every month, all 33 models, no card.