Skip to main content

Error format

Failed requests return a JSON body with an error object:
Some codes add extra fields, listed below. New codes may be added over time, so handle unknown codes by HTTP status. OpenAI and Anthropic SDKs choose their exception class from the HTTP status. On /v1/messages and /v1/messages/count_tokens, errors use Anthropic’s wrapper instead: {"type": "error", "error": {"type": ..., "message": ..., "code": ..., "request_id": ...}}, with the same codes and extra fields. See Messages errors.

Request IDs

Responses carry an x-request-id header with Wafer’s ID for the request, a 12-character hex string. Log it, and include it when you contact support. If you send your own x-request-id header, Wafer records it with the request, but the response header carries Wafer’s ID. A few 429 responses from Wafer’s front proxy have no request ID.

Status codes

Wafer usually returns transient server-side failures as 429 with a Retry-After header rather than 5xx, so SDK retry logic for rate limits covers them. Treat any 5xx you do receive as retryable too.

Error codes

Authentication and access

  • missing_api_key (401): No Authorization: Bearer <key> or x-api-key header.
  • invalid_api_key (401): The key doesn’t exist or was revoked.
  • model_not_allowed (403): The model exists, but your key’s type or account can’t use it.
  • model_not_found (404): Unknown model, or a model your account doesn’t have access to. The message lists the available models; GET /v1/models lists the models visible to your key.
  • not_found (404): The path isn’t a Wafer Serverless route.

Request validation (400)

  • invalid_json_request: The request body isn’t valid JSON, or isn’t a JSON object.
  • missing_model: The request has no model.
  • context_length_exceeded: The request doesn’t fit the model’s context window. Extra fields: model, context_length_limit, and sometimes suggested_models (models with larger windows). Shorten the prompt or switch models.
  • unsupported_value: A field has a value the model doesn’t accept, such as an unknown reasoning_effort. param names the field.
  • unsupported_feature: The model doesn’t support a requested feature, such as n > 1. param names the field.
  • unsupported_parameter: The API doesn’t support the field, such as previous_response_id on /v1/responses.
  • unsupported_response_format: The model doesn’t support the response_format type.
  • unsupported_regex: The model doesn’t support top-level regex.
  • json_schema_refs_unsupported, json_schema_refs_unresolved, json_schema_refs_recursive, json_schema_refs_too_deep, json_schema_refs_too_large: A schema has $ref references that can’t be inlined. Inline them yourself or simplify the schema.
  • tool_schema_invalid: A tool’s parameters isn’t a JSON Schema object with "type": "object".
  • missing_tool_name, duplicate_tool_name: A tool has no name, or two tools share a name.
  • tool_choice_unknown_tool: tool_choice names a tool that isn’t in tools.
  • missing_tool_call_id, orphan_tool_message: A tool message has no tool_call_id, or its ID doesn’t match a tool call in an earlier assistant message.
  • unsupported_input_item, unsupported_content_type, unsupported_block_type: An input item, content part, or content block type isn’t supported on this API.
  • invalid_zdr_header: Wafer-ZDR is set to a value other than required.
  • model_request_rejected: The model rejected the request, for example because an image couldn’t be fetched or was sent to a model without vision. Check the request before retrying. This code keeps the status the model returned, usually 400.
Other invalid_*, missing_*, and empty_* codes describe a malformed field; param or message identifies it.

Credits (402)

  • insufficient_credits: Your balance is below the estimated cost of the request. Extra fields: credits_available_cents and credits_required_cents_estimate. Add credits at app.wafer.ai, then retry.
  • member_spend_limit: A spend limit set on your account for this member has been reached.
The credit check happens before the request runs. The estimate assumes the request generates its full max_tokens (times n). If you omit max_tokens, the estimate uses the default (see Output length), which can be tens of cents to over a dollar per request. With a low balance, set a smaller max_tokens.

ZDR (422)

  • model_zdr_not_supported: ZDR was required by the Wafer-ZDR: required header or by account policy, but the model doesn’t support ZDR. See Zero Data Retention.

Rate limits and capacity (429)

  • server_overloaded: The model is temporarily at capacity, or a transient failure happened while serving the request. Retry after Retry-After.
  • concurrency_limit_exceeded: Your account has too many requests in flight. Retry when an earlier request finishes.
  • model_service_unavailable, model_request_timeout, auth_backend_unavailable, edge_at_capacity: Transient failures. Retry after Retry-After.
  • rate_limited: An unexpected error on Wafer’s side. Retry after Retry-After; if it repeats, contact support with the request ID.

Rate limits

Wafer doesn’t publish fixed requests-per-minute or tokens-per-minute limits. Besides transient failures, two things return 429:
  • Model capacity. When a model is busy, Wafer sheds new requests with server_overloaded instead of queueing them for a long time. This protects latency for requests already running and usually clears within seconds.
  • Account concurrency. An account can have a limit on concurrent in-flight requests. Exceeding it returns concurrency_limit_exceeded.
429 responses include Retry-After (seconds). Capacity and concurrency 429s generated by Wafer’s API also include RateLimit-Reset with the same value. Successful responses carry no rate-limit headers. If you need guaranteed throughput, contact support@wafer.ai.

Retries

  • Retry 429 after the Retry-After delay, using exponential backoff with jitter and a cap on attempts. The OpenAI and Anthropic SDKs do this by default.
  • Don’t retry other 4xx errors unchanged. Fix the request first. For 402, add credits first.
  • Wafer already retries some failures internally before sending the first byte, so a 429 means those retries didn’t succeed.
  • Requests that fail before generation starts are not charged. A stream that fails after it starts may be charged for tokens already generated.

Errors during streaming

Once a stream has started, the HTTP status is already 200, so errors arrive in the stream or as a dropped connection.
  • Chat Completions: an error arrives as a data: line with an error object instead of a chunk, for example data: {"error": {"message": "...", "type": "rate_limit_error", "param": null, "code": "server_overloaded"}}. It can still be followed by data: [DONE], so [DONE] alone doesn’t mean success. Treat any error line as a failed request.
  • Messages: an event: error event carries {"type": "error", "error": {"type": "overloaded_error" or "api_error", "message": ..., "code": ...}}, and the stream ends without message_stop.
  • Responses: a response.failed event carries the response with "status": "failed" and error: {"code": ..., "message": ...}, and the stream ends without response.completed.
  • The code in these events can be one of the codes in Rate limits and capacity, model_request_rejected, or invalid_stream_response (the model returned a malformed stream). server_overloaded (overloaded_error on Messages) and invalid_stream_response are safe to retry.
  • All APIs: a stream that ends without its final event (data: [DONE], message_stop, or response.completed) failed; retry the request.

Timeouts

  • A non-streaming request that runs longer than about 15 minutes (longer on some models) can fail with server_overloaded.
  • A stream can drop, without an error event, if no data arrives for about 15 minutes.
Use streaming for long generations. For non-streaming requests with large reasoning outputs, set your client timeout to at least 15 minutes.

Request example

This request fails with unsupported_value because reasoning_effort isn’t a supported value: