Skip to content

Messages API (Anthropic wire dialect) ​

/api/client/v1/messages speaks Anthropic's Messages API request/response shape natively, alongside the OpenAI-shaped Chat Completions endpoint. Both run through the same model routing, guardrails, and provider gateway — pick whichever wire format your client already speaks.

New in September 2026

This endpoint shipped alongside the guardrail hook plane rewrite. If your integration currently targets Anthropic's API directly, pointing it at this endpoint instead gets you Console's model routing, guardrails, and tracing without a rewrite.

Endpoint ​

POST /api/client/v1/messages
Authorization: Bearer <token>
Content-Type: application/json

Request ​

json
{
  "model": "claude-sonnet",
  "system": "You are a helpful assistant.",
  "messages": [
    { "role": "user", "content": "What is the capital of France?" }
  ],
  "max_tokens": 1000,
  "temperature": 0.7,
  "stream": false
}
FieldTypeRequiredDescription
modelstringYesModel key configured in the dashboard
messagesarrayYesArray of { role: "user" | "assistant", content } turns
systemstringNoSystem prompt, kept separate from messages as Anthropic's shape expects
max_tokensnumberYesMaximum tokens to generate
temperaturenumberNoSampling temperature (0-1)
streambooleanNoEnable SSE streaming (default: false)
toolsarrayNoTool definitions, in Anthropic's tool-use shape

content accepts either a plain string or an array of content blocks ({ type: "text", text: "..." }, plus tool-use/tool-result blocks when tools is set).

Response (Non-Streaming) ​

json
{
  "id": "msg_abc123",
  "type": "message",
  "role": "assistant",
  "model": "claude-sonnet",
  "content": [
    { "type": "text", "text": "The capital of France is Paris." }
  ],
  "stop_reason": "end_turn",
  "usage": {
    "input_tokens": 25,
    "output_tokens": 8
  }
}

Response (Streaming) ​

When stream: true, the response is a Server-Sent Events stream using Anthropic's event types (message_start, content_block_start, content_block_delta, content_block_stop, message_delta, message_stop) rather than Chat Completions' chat.completion.chunk shape.

Guardrails ​

Guardrails attached to the target model apply the same way they do on Chat Completions — the input.pre / output.pre hooks run regardless of which wire dialect the request arrived through. A blocked request returns the same guardrail_block error shape.

Errors ​

StatusDescription
400Missing model, messages, or max_tokens; guardrail block
401Invalid API token
500Model not found or not configured for chat completions
429Rate limit or quota exceeded

Examples ​

cURL ​

bash
curl -X POST https://gateway.example.com/api/client/v1/messages \
  -H "Authorization: Bearer cpeer_your_token" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "claude-sonnet",
    "max_tokens": 1000,
    "messages": [{"role": "user", "content": "Hello"}]
  }'

Python (Anthropic SDK) ​

Point Anthropic's own SDK at the gateway by overriding its base URL:

python
from anthropic import Anthropic

client = Anthropic(
    api_key="cpeer_your_token",
    base_url="https://gateway.example.com/api/client/v1",
)
response = client.messages.create(
    model="claude-sonnet",
    max_tokens=1000,
    messages=[{"role": "user", "content": "Hello"}],
)

Studio · Pulse · Console · Agent SDK and more — the Cognipeer documentation hub