OpenAI vs Anthropic API format: what's actually different?
Chat Completions, Responses and Messages side by side: endpoints, auth headers, system prompts, streaming events and where token usage lives.
Most AI tools speak one of two dialects: OpenAI's (Chat Completions, and the newer Responses API) or Anthropic's (Messages). Apixo accepts both, for every model — but knowing the differences helps when you configure tools or debug requests.
Endpoints and auth
| OpenAI | Anthropic | |
|---|---|---|
| Endpoint | /v1/chat/completions, /v1/responses |
/v1/messages |
| Auth header | Authorization: Bearer <key> |
x-api-key: <key> |
| Version header | — | anthropic-version: 2023-06-01 |
| Base URL in SDKs | usually ends with /v1 |
usually without /v1 |
On Apixo either auth header works everywhere.
Request shape
OpenAI puts the system prompt in the messages array:
{
"model": "gpt-6-luna",
"messages": [
{"role": "system", "content": "You are terse."},
{"role": "user", "content": "Hi"}
]
}Anthropic uses a top-level system field and requires max_tokens:
{
"model": "claude-sonnet-5",
"max_tokens": 1024,
"system": "You are terse.",
"messages": [{"role": "user", "content": "Hi"}]
}Streaming
Both use Server-Sent Events, with different event shapes:
- Chat Completions:
data: {...chunk...}lines ending withdata: [DONE]. Usage arrives near the end whenstream_options.include_usageis set (we enable it for you). - Anthropic: named events —
message_start,content_block_delta,message_delta,message_stop. Input usage is inmessage_start, output usage inmessage_delta. - Responses: typed events such as
response.output_text.delta; final usage is inresponse.completed.
Token usage fields
| Input | Output | Cached input | |
|---|---|---|---|
| Chat Completions | prompt_tokens (includes cached) |
completion_tokens |
prompt_tokens_details.cached_tokens |
| Anthropic | input_tokens (excludes cached) |
output_tokens |
cache_read_input_tokens |
| Responses | input_tokens (includes cached) |
output_tokens |
input_tokens_details.cached_tokens |
Watch out for the includes vs excludes difference when you compute costs yourself. Our dashboard normalises all three so every request shows uncached input, cached input and output separately.
Which should you use?
Use whatever your tool expects. Anthropic-native tools (Claude Code, the Anthropic SDK) → Messages. Almost everything else → Chat Completions. The model you call doesn't need to match the format: on Apixo you can call Claude through Chat Completions and GPT through Messages.
One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.
Get your API key