Error format
Failed requests return a JSON body with anerror object:
Some codes add extra fields, listed below. New codes may be added over time, so handle unknown codes by HTTP status. OpenAI and Anthropic SDKs choose their exception class from the HTTP status.
On
/v1/messages and /v1/messages/count_tokens, errors use Anthropic’s wrapper instead: {"type": "error", "error": {"type": ..., "message": ..., "code": ..., "request_id": ...}}, with the same codes and extra fields. See Messages errors.
Request IDs
Responses carry anx-request-id header with Wafer’s ID for the request, a 12-character hex string. Log it, and include it when you contact support. If you send your own x-request-id header, Wafer records it with the request, but the response header carries Wafer’s ID. A few 429 responses from Wafer’s front proxy have no request ID.
Status codes
Wafer usually returns transient server-side failures as
429 with a Retry-After header rather than 5xx, so SDK retry logic for rate limits covers them. Treat any 5xx you do receive as retryable too.
Error codes
Authentication and access
missing_api_key(401): NoAuthorization: Bearer <key>orx-api-keyheader.invalid_api_key(401): The key doesn’t exist or was revoked.model_not_allowed(403): The model exists, but your key’s type or account can’t use it.model_not_found(404): Unknownmodel, or a model your account doesn’t have access to. The message lists the available models;GET /v1/modelslists the models visible to your key.not_found(404): The path isn’t a Wafer Serverless route.
Request validation (400)
invalid_json_request: The request body isn’t valid JSON, or isn’t a JSON object.missing_model: The request has nomodel.context_length_exceeded: The request doesn’t fit the model’s context window. Extra fields:model,context_length_limit, and sometimessuggested_models(models with larger windows). Shorten the prompt or switch models.unsupported_value: A field has a value the model doesn’t accept, such as an unknownreasoning_effort.paramnames the field.unsupported_feature: The model doesn’t support a requested feature, such asn > 1.paramnames the field.unsupported_parameter: The API doesn’t support the field, such asprevious_response_idon/v1/responses.unsupported_response_format: The model doesn’t support theresponse_formattype.unsupported_regex: The model doesn’t support top-levelregex.json_schema_refs_unsupported,json_schema_refs_unresolved,json_schema_refs_recursive,json_schema_refs_too_deep,json_schema_refs_too_large: A schema has$refreferences that can’t be inlined. Inline them yourself or simplify the schema.tool_schema_invalid: A tool’sparametersisn’t a JSON Schema object with"type": "object".missing_tool_name,duplicate_tool_name: A tool has no name, or two tools share a name.tool_choice_unknown_tool:tool_choicenames a tool that isn’t intools.missing_tool_call_id,orphan_tool_message: Atoolmessage has notool_call_id, or its ID doesn’t match a tool call in an earlier assistant message.unsupported_input_item,unsupported_content_type,unsupported_block_type: An input item, content part, or content block type isn’t supported on this API.invalid_zdr_header:Wafer-ZDRis set to a value other thanrequired.model_request_rejected: The model rejected the request, for example because an image couldn’t be fetched or was sent to a model without vision. Check the request before retrying. This code keeps the status the model returned, usually400.
invalid_*, missing_*, and empty_* codes describe a malformed field; param or message identifies it.
Credits (402)
insufficient_credits: Your balance is below the estimated cost of the request. Extra fields:credits_available_centsandcredits_required_cents_estimate. Add credits at app.wafer.ai, then retry.member_spend_limit: A spend limit set on your account for this member has been reached.
max_tokens (times n). If you omit max_tokens, the estimate uses the default (see Output length), which can be tens of cents to over a dollar per request. With a low balance, set a smaller max_tokens.
ZDR (422)
model_zdr_not_supported: ZDR was required by theWafer-ZDR: requiredheader or by account policy, but the model doesn’t support ZDR. See Zero Data Retention.
Rate limits and capacity (429)
server_overloaded: The model is temporarily at capacity, or a transient failure happened while serving the request. Retry afterRetry-After.concurrency_limit_exceeded: Your account has too many requests in flight. Retry when an earlier request finishes.model_service_unavailable,model_request_timeout,auth_backend_unavailable,edge_at_capacity: Transient failures. Retry afterRetry-After.rate_limited: An unexpected error on Wafer’s side. Retry afterRetry-After; if it repeats, contact support with the request ID.
Rate limits
Wafer doesn’t publish fixed requests-per-minute or tokens-per-minute limits. Besides transient failures, two things return429:
- Model capacity. When a model is busy, Wafer sheds new requests with
server_overloadedinstead of queueing them for a long time. This protects latency for requests already running and usually clears within seconds. - Account concurrency. An account can have a limit on concurrent in-flight requests. Exceeding it returns
concurrency_limit_exceeded.
429 responses include Retry-After (seconds). Capacity and concurrency 429s generated by Wafer’s API also include RateLimit-Reset with the same value. Successful responses carry no rate-limit headers.
If you need guaranteed throughput, contact support@wafer.ai.
Retries
- Retry
429after theRetry-Afterdelay, using exponential backoff with jitter and a cap on attempts. The OpenAI and Anthropic SDKs do this by default. - Don’t retry other
4xxerrors unchanged. Fix the request first. For402, add credits first. - Wafer already retries some failures internally before sending the first byte, so a
429means those retries didn’t succeed. - Requests that fail before generation starts are not charged. A stream that fails after it starts may be charged for tokens already generated.
Errors during streaming
Once a stream has started, the HTTP status is already200, so errors arrive in the stream or as a dropped connection.
- Chat Completions: an error arrives as a
data:line with anerrorobject instead of a chunk, for exampledata: {"error": {"message": "...", "type": "rate_limit_error", "param": null, "code": "server_overloaded"}}. It can still be followed bydata: [DONE], so[DONE]alone doesn’t mean success. Treat anyerrorline as a failed request. - Messages: an
event: errorevent carries{"type": "error", "error": {"type": "overloaded_error" or "api_error", "message": ..., "code": ...}}, and the stream ends withoutmessage_stop. - Responses: a
response.failedevent carries the response with"status": "failed"anderror: {"code": ..., "message": ...}, and the stream ends withoutresponse.completed. - The
codein these events can be one of the codes in Rate limits and capacity,model_request_rejected, orinvalid_stream_response(the model returned a malformed stream).server_overloaded(overloaded_erroron Messages) andinvalid_stream_responseare safe to retry. - All APIs: a stream that ends without its final event (
data: [DONE],message_stop, orresponse.completed) failed; retry the request.
Timeouts
- A non-streaming request that runs longer than about 15 minutes (longer on some models) can fail with
server_overloaded. - A stream can drop, without an error event, if no data arrives for about 15 minutes.
Request example
This request fails withunsupported_value because reasoning_effort isn’t a supported value: