Messages API (Anthropic wire dialect)
/api/client/v1/messages speaks Anthropic's Messages API request/response shape natively, alongside the OpenAI-shaped Chat Completions endpoint. Both run through the same model routing, guardrails, and provider gateway — pick whichever wire format your client already speaks.
New in September 2026
This endpoint shipped alongside the guardrail hook plane rewrite. If your integration currently targets Anthropic's API directly, pointing it at this endpoint instead gets you Console's model routing, guardrails, and tracing without a rewrite.
Endpoint
POST /api/client/v1/messages
Authorization: Bearer <token>
Content-Type: application/json2
3
Request
{
"model": "claude-sonnet",
"system": "You are a helpful assistant.",
"messages": [
{ "role": "user", "content": "What is the capital of France?" }
],
"max_tokens": 1000,
"temperature": 0.7,
"stream": false
}2
3
4
5
6
7
8
9
10
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model key configured in the dashboard |
messages | array | Yes | Array of { role: "user" | "assistant", content } turns |
system | string | No | System prompt, kept separate from messages as Anthropic's shape expects |
max_tokens | number | Yes | Maximum tokens to generate |
temperature | number | No | Sampling temperature (0-1) |
stream | boolean | No | Enable SSE streaming (default: false) |
tools | array | No | Tool definitions, in Anthropic's tool-use shape |
content accepts either a plain string or an array of content blocks ({ type: "text", text: "..." }, plus tool-use/tool-result blocks when tools is set).
Response (Non-Streaming)
{
"id": "msg_abc123",
"type": "message",
"role": "assistant",
"model": "claude-sonnet",
"content": [
{ "type": "text", "text": "The capital of France is Paris." }
],
"stop_reason": "end_turn",
"usage": {
"input_tokens": 25,
"output_tokens": 8
}
}2
3
4
5
6
7
8
9
10
11
12
13
14
Response (Streaming)
When stream: true, the response is a Server-Sent Events stream using Anthropic's event types (message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop) rather than Chat Completions' chat.completion.chunk shape.
Guardrails
Guardrails attached to the target model apply the same way they do on Chat Completions — the input.pre / output.pre hooks run regardless of which wire dialect the request arrived through. A blocked request returns the same guardrail_block error shape.
Errors
| Status | Description |
|---|---|
| 400 | Missing model, messages, or max_tokens; guardrail block |
| 401 | Invalid API token |
| 500 | Model not found or not configured for chat completions |
| 429 | Rate limit or quota exceeded |
Examples
cURL
curl -X POST https://gateway.example.com/api/client/v1/messages \
-H "Authorization: Bearer cpeer_your_token" \
-H "Content-Type: application/json" \
-d '{
"model": "claude-sonnet",
"max_tokens": 1000,
"messages": [{"role": "user", "content": "Hello"}]
}'2
3
4
5
6
7
8
Python (Anthropic SDK)
Point Anthropic's own SDK at the gateway by overriding its base URL:
from anthropic import Anthropic
client = Anthropic(
api_key="cpeer_your_token",
base_url="https://gateway.example.com/api/client/v1",
)
response = client.messages.create(
model="claude-sonnet",
max_tokens=1000,
messages=[{"role": "user", "content": "Hello"}],
)2
3
4
5
6
7
8
9
10
11

