Skip to content
Apixo
Blog
guide· 2 min read

OpenAI vs Anthropic API format: what's actually different?

Chat Completions, Responses and Messages side by side: endpoints, auth headers, system prompts, streaming events and where token usage lives.

Most AI tools speak one of two dialects: OpenAI's (Chat Completions, and the newer Responses API) or Anthropic's (Messages). Apixo accepts both, for every model — but knowing the differences helps when you configure tools or debug requests.

Endpoints and auth

OpenAI Anthropic
Endpoint /v1/chat/completions, /v1/responses /v1/messages
Auth header Authorization: Bearer <key> x-api-key: <key>
Version header — anthropic-version: 2023-06-01
Base URL in SDKs usually ends with /v1 usually without /v1

On Apixo either auth header works everywhere.

Request shape

OpenAI puts the system prompt in the messages array:

{
  "model": "gpt-6-luna",
  "messages": [
    {"role": "system", "content": "You are terse."},
    {"role": "user", "content": "Hi"}
  ]
}

Anthropic uses a top-level system field and requires max_tokens:

{
  "model": "claude-sonnet-5",
  "max_tokens": 1024,
  "system": "You are terse.",
  "messages": [{"role": "user", "content": "Hi"}]
}

Streaming

Both use Server-Sent Events, with different event shapes:

  • Chat Completions: data: {...chunk...} lines ending with data: [DONE]. Usage arrives near the end when stream_options.include_usage is set (we enable it for you).
  • Anthropic: named events — message_start, content_block_delta, message_delta, message_stop. Input usage is in message_start, output usage in message_delta.
  • Responses: typed events such as response.output_text.delta; final usage is in response.completed.

Token usage fields

Input Output Cached input
Chat Completions prompt_tokens (includes cached) completion_tokens prompt_tokens_details.cached_tokens
Anthropic input_tokens (excludes cached) output_tokens cache_read_input_tokens
Responses input_tokens (includes cached) output_tokens input_tokens_details.cached_tokens

Watch out for the includes vs excludes difference when you compute costs yourself. Our dashboard normalises all three so every request shows uncached input, cached input and output separately.

Which should you use?

Use whatever your tool expects. Anthropic-native tools (Claude Code, the Anthropic SDK) → Messages. Almost everything else → Chat Completions. The model you call doesn't need to match the format: on Apixo you can call Claude through Chat Completions and GPT through Messages.

#api#openai#anthropic
Try it with your own tools

One key for Claude, GPT, GLM, DeepSeek and more. Pay per token with crypto.

Get your API key

Keep reading