Vision (Image Input)

Send images alongside text for analysis, OCR, and UI understanding.

On this page

Vision#

Models advertising the Vision capability on the Models page accept images as input. Instead of a plain string, content becomes an array of parts.

Base64 image#

{
  "model": "claude-opus-4.8",
  "messages": [{
    "role": "user",
    "content": [
      {"type": "text", "text": "What is in this image?"},
      {
        "type": "image_url",
        "image_url": {"url": "data:image/png;base64,iVBORw0KGgoAAAANS..."}
      }
    ]
  }]
}

Image by URL#

{
  "type": "image_url",
  "image_url": {"url": "https://example.com/screenshot.png"}
}

Public URLs must be reachable from our servers. Base64 is more reliable for local or private files.

Multiple images#

Add as many image_url parts as you need in one message:

{
  "role": "user",
  "content": [
    {"type": "text", "text": "What changed between these two screenshots?"},
    {"type": "image_url", "image_url": {"url": "data:image/png;base64,AAA..."}},
    {"type": "image_url", "image_url": {"url": "data:image/png;base64,BBB..."}}
  ]
}

Supported formats and limits#

Property Value
Formats PNG, JPEG, WebP, GIF (first frame)
Max per image 5 MB
Max request payload 8 MB total

Python example#

import base64
from openai import OpenAI

client = OpenAI(api_key="cc_your_key_here", base_url="https://cleanapis.com/v1")

with open("screenshot.png", "rb") as f:
    encoded = base64.b64encode(f.read()).decode()

response = client.chat.completions.create(
    model="claude-opus-4.8",
    messages=[{
        "role": "user",
        "content": [
            {"type": "text", "text": "Describe this UI and list its buttons."},
            {"type": "image_url", "image_url": {
                "url": f"data:image/png;base64,{encoded}"
            }},
        ],
    }],
)
print(response.choices[0].message.content)

Billing#

Images consume prompt tokens based on their resolution. A typical screenshot costs a few hundred to a few thousand prompt tokens, reported in usage.prompt_tokens as usual.

Errors#

Sending an image to a model without the Vision capability returns 422:

{
  "error": {
    "message": "This model does not support image input.",
    "type": "invalid_request_error",
    "code": "unsupported_content"
  }
}

Check capabilities in GET /models before sending images.

Still stuck?

Open a support ticket from your dashboard and we'll take a look.

Get help