Skip to main content
POST https://pass.wafer.ai/v1/messages accepts Anthropic Messages API requests. Use it with the Anthropic SDKs, Claude Code, and other Anthropic-format clients. Set the client’s base URL to https://pass.wafer.ai; the client appends /v1/messages.

Request

Authenticate with x-api-key or Authorization: Bearer. The anthropic-version and anthropic-beta headers are accepted and ignored. Use a Wafer model ID from GET /v1/models as model. Anthropic model names such as claude-... are not available.

Supported fields

metadata and cache_control are accepted and ignored. Prompt caching is automatic; see Prompt caching.

Response

stop_reason is end_turn, max_tokens, stop_sequence, or tool_use.
usage.input_tokens counts all prompt tokens, including cache_read_input_tokens. This differs from Anthropic’s API. See Usage and Billing.

Thinking

On /v1/messages, reasoning is off unless you enable it, whatever the model’s default on other APIs.
  • thinking: {"type": "enabled"} turns reasoning on. Set output_config.effort to choose the effort; without it, the effort depends on the model.
  • budget_tokens is accepted but ignored. Use output_config.effort and max_tokens to control reasoning length.
  • Reasoning is returned as thinking content blocks before the answer. Their signature is an empty string.
  • You can pass earlier thinking blocks back in assistant turns.
  • On GLM-5.3 and GLM-5.3-Flash, reasoning can’t be fully turned off. With thinking off, the model can still spend output tokens reasoning that isn’t returned; they count toward max_tokens and usage.output_tokens. See the reasoning warning.

Tools

When the model calls a tool, the response contains a tool_use block and stop_reason is tool_use. Send the result back in a tool_result block:

Vision

On models whose catalog card has wafer.capabilities.messages.vision: true, send images as base64 image blocks:
url image sources are also accepted; Wafer fetches them, and an image that can’t be fetched fails the request with 400 code model_request_rejected. Base64 avoids fetch failures. Images sent to models without vision support fail with 400 code model_request_rejected.

Count tokens

POST /v1/messages/count_tokens takes the same body without max_tokens and returns the prompt token count:
If the exact count isn’t available, the response is an estimate and carries the header x-wafer-input-tokens-estimated: true.

Streaming

With "stream": true, Wafer sends Anthropic-format server-sent events: message_start, then content_block_start, content_block_delta, and content_block_stop for each content block, then message_delta with stop_reason and final usage, then message_stop. Text arrives as text_delta deltas. Reasoning arrives in thinking blocks as thinking_delta deltas when thinking is on. Tool input arrives as input_json_delta deltas; concatenate partial_json to get the full input. Wafer doesn’t send ping or signature_delta events.
message_start carries zero usage. Read final usage from message_delta. If generation fails after the stream starts, Wafer sends an error event and ends the stream without message_stop. If the connection closes before message_stop, the request failed too. See Errors during streaming.

Errors

Errors on /v1/messages and /v1/messages/count_tokens use Anthropic’s error format, with Wafer’s code and request_id added:
Anthropic SDKs pick the exception class from the HTTP status. In your own code, branch on the HTTP status, then on error.code. See Errors and Rate Limits for the codes.