Skip to content
Apixo
Reference

API reference

Endpoints, authentication, streaming, errors and rate limits of the Apixo API.

Authentication

Send your key in either header — both work on every endpoint:

Authorization: Bearer sk-apixo-...
x-api-key: sk-apixo-...

Endpoints

Method Path Format
POST /v1/chat/completions OpenAI Chat Completions
POST /v1/responses OpenAI Responses
POST /v1/messages Anthropic Messages
POST /v1/messages/count_tokens Anthropic token counting (not billed)
GET /v1/models List of models available to you

Every model works on every format — you can call claude-fable-5 through /v1/chat/completions or gpt-6-astra through /v1/messages. Request bodies are forwarded unchanged, so tools, images, reasoning / thinking parameters and JSON mode work as documented by OpenAI and Anthropic.

Streaming

Set "stream": true. Responses are sent as Server-Sent Events exactly like the original APIs. For Chat Completions we automatically enable stream_options.include_usage, so the last chunk contains token usage.

Usage & billing

Each response includes the standard usage object. We bill:

  • input tokens at the model's input price,
  • output tokens (including reasoning) at the output price,
  • cached input tokens at the lower cache price,
  • cache writes (Anthropic cache_creation_input_tokens) at the cache-write price.

Prices are per 1M tokens — see pricing. The cost of each request appears in your usage log within a second.

Errors

Errors use the format of the endpoint you called (OpenAI {"error": {...}} or Anthropic {"type": "error", ...}).

Status Meaning What to do
400 Invalid JSON or the model isn't available Check the model ID against /v1/models
401 Missing, wrong or disabled key Copy the key again from the dashboard
402 Balance too low Top up
403 Account suspended, or a trial account calling a non-trial model Top up or contact support
429 Rate limit reached Wait for the retry-after seconds
5xx Temporary upstream problem Retry with backoff; check status

Rate limits

Limits are per API key, in requests per minute. Paying accounts get a higher limit than trial accounts. When you hit it you receive 429 with a retry-after header.

Privacy

We never store prompt or response content. For billing we keep the model, token counts, cost, latency and status of each request.