> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://www.respan.ai/docs/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://www.respan.ai/docs/_mcp/server.

# Create a chat completion

POST https://api.respan.ai/api/chat/completions
Content-Type: application/json

Sends a chat completion request through the Respan gateway with automatic logging. Accepts [OpenAI chat completion parameters](https://platform.openai.com/docs/apis/chat) and Respan options for fallbacks, caching, and prompt management.

Pass Respan parameters in top-level body fields, under `respan_params`, or as base64-encoded JSON in the `X-Data-Respan-Params` header. If the same field is sent more than one way, `respan_params` takes precedence over the header, and both take precedence over a top-level field. The exception is `variables`: top-level `variables` are merged key by key over the `variables` from `respan_params` or the header. `respan_params`, `keywordsai_params`, and the decoded header must each be a JSON object; a top-level `respan_params` may also be a JSON string that decodes to one. Any other value is ignored, and so are the header's parameters: the request is served, but none of those parameters apply. Fields sent at the top level still apply. With the OpenAI SDK, use `extra_body`.

For legacy compatibility, `keywordsai_params` is merged into `respan_params`, and `X-Data-Keywordsai-Params` is still accepted and renamed internally.

Reference: https://www.respan.ai/docs/apis/gateway/create-chat-completion

## Authentication

- `Authorization` header (bearer token, required) — Use your Respan API key for Respan API authentication. Enter only the Respan API key value; clients send Authorization: Bearer \<RESPAN\_API\_KEY>. For /api/responses, provider credentials such as Perplexity, OpenAI, or Azure OpenAI go in Settings -> Providers or respan\_params.credential\_override in the request body, not in this authentication field.

## Request

### Headers

- `X-Data-Respan-Params` (string, optional) — Base64-encoded JSON object of Respan parameters. Legacy `X-Data-Keywordsai-Params` is still accepted.
- `X-Respan-Route-Provider` (string, optional) — Pin the request to a specific provider without changing the model slug. Example: `vertex_ai` routes a `claude-sonnet-4-5-20250929` request to Vertex AI Claude.
- `X-Respan-Beta` (string, optional) — Comma-separated beta feature flags. Available: token-breakdown-2026-03-26, env-scoped-integrations-2026-03-28

### Body (application/json)

This endpoint expects an object.

- `messages` (list of ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMessagesItems, required) — Array of messages in the conversation. Each message has `role` (`system`, `user`, `assistant`, `tool`) and `content`.
- `model` (string, required) — Model to use. See [Models](https://platform.respan.ai/platform/models) for available options.
- `stream` (boolean, optional) — Stream back partial progress token by token as server-sent events.
- `tools` (list of ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaToolsItems, optional) — Tools the model may call. Currently only functions are supported.
- `tool_choice` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaToolChoice, optional) — Controls tool selection. `"none"` = no tools, `"auto"` = model decides, or specify a tool object.
- `frequency_penalty` (double, optional) — Penalizes tokens based on frequency in text so far (-2 to 2).
- `max_tokens` (double, optional) — Maximum tokens to generate.
- `temperature` (double, optional, default: 1) — Sampling temperature (0-2). Higher = more random.
- `n` (double, optional, default: 1) — Number of completions to generate. Note: costs multiply with `n`.
- `logprobs` (boolean, optional) — Return log probabilities of output tokens.
- `echo` (boolean, optional) — Echo back the prompt in addition to the completion
- `stop` (list of string, optional) — Stop sequences where generation halts.
- `presence_penalty` (double, optional) — Penalizes tokens already present in text (-2 to 2).
- `logit_bias` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLogitBias, optional) — Used to modify the probability of tokens appearing in the response
- `response_format` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaResponseFormat, optional) — Output format. Set `{"type": "json_schema", "json_schema": {...}}` for structured output, or `{"type": "json_object"}` for JSON mode.
- `parallel_tool_calls` (boolean, optional) — Enable parallel function calling during tool use.
- `load_balance_group` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLoadBalanceGroup, optional) — Load balance group selection. Use `{"group_id": "..."}` to route through a configured group.
- `fallback_models` (list of string, optional) — Backup models (ranked by priority) if the primary model fails.
- `customer_credentials` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCustomerCredentials, optional) — Per-customer LLM provider credentials. Keys are provider names, values are API keys.
- `credential_override` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCredentialOverride, optional) — One-off credential overrides per provider. Overrides uploaded provider keys for this request only.
- `cache_enabled` (boolean, optional) — Enable response caching. See [Caching](/docs/documentation/features/gateway/caching).
- `cache_ttl` (double, optional) — How long a cached response is served, in seconds. A response stored without `cache_ttl` is served from the cache for at most 30 minutes.
- `cache_options` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCacheOptions, optional) — Cache behavior options. Properties: `cache_by_customer`, `is_cached_by_model`, `omit_log`.
- `prompt` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaPrompt, optional) — Prompt template config. Properties: `prompt_id` (required), `variables` (template variables), `version` (number, or `"latest"` for draft), `echo` (return rendered prompt), `override` (use override_params), `override_params` (OpenAI params to override), `schema_version` (`1` = legacy, `2` = prompt config wins). See [Prompt management](/docs/documentation/features/prompt-management/advanced).
- `retry_params` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaRetryParams, optional) — Has no effect on this endpoint: chat completions always uses your organization's retry settings. Per-request `retry_params` (`retry_enabled`, `num_retries`, `retry_after`) applies only to `POST /api/responses`. See [Retries and fallback](/docs/documentation/features/gateway/retries).
- `disable_log` (boolean, optional) — When `true`, omits input/output from the log. Metrics (tokens, cost, latency) are still recorded.
- `model_name_map` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaModelNameMap, optional) — Azure deployment name mapping. Maps your custom Azure deployment names to standard model names.
- `models` (list of string, optional) — Model list for LLM router selection.
- `exclude_providers` (list of string, optional) — Providers to exclude from routing. All models under excluded providers are skipped.
- `exclude_models` (list of string, optional) — Specific models to exclude from routing.
- `metadata` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMetadata, optional) — Custom key-value metadata attached to the span.
- `custom_identifier` (string, optional) — Indexed custom tag for fast querying.
- `customer_identifier` (string, optional) — End user identifier for analytics and budgets.
- `customer_params` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCustomerParams, optional) — Customer details. Properties: `customer_identifier` (takes precedence over the top-level `customer_identifier`), `name` and `email` (logged with the request, and saved on the customer when Respan first sees it), and `rate_limit` (requests per minute for this customer, overriding your organization's customer rate limit; requests over it get `429`). Budget fields sent here aren't saved or enforced. Set budgets with [Update a user](/docs/apis/users/update-user).
- `request_breakdown` (boolean, optional) — Return response metrics summary in the response body. For streaming, metrics appear in the final chunk.
- `positive_feedback` (boolean, optional) — User feedback. `true` = liked, `false` = disliked.
- `load_balance_models` (list of ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLoadBalanceModelsItems, optional) — Inline load balancing options. Each item can include `model`, `weight`, and optional `credentials`.
- `thread_identifier` (string, optional) — Conversation thread ID. Spans with the same `thread_identifier` are grouped together.
- `properties` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchema, optional) — Typed metadata preserving native types (numbers, booleans, nested objects). Unlike `metadata` which coerces to strings.
- `retries` (integer, optional, default: 0) — Has no effect: chat completions always uses your organization's retry settings. See [Retries and fallback](/docs/documentation/features/gateway/retries).
- `weight` (double, optional) — Load balancing weight.
- `span_name` (string, optional) — Custom span name for tracing.
- `respan_params` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaRespanParams, optional) — Namespaced container for all Respan parameters. Alternative to passing them at top level. Must be a JSON object, or a JSON string that decodes to one. A list, number or boolean is ignored and none of its params apply.

## Response

### 200

Successful response for Create chat completion

- `id` (string, required) — Chat completion ID.
- `object` (string, required)
- `created` (integer, required) — Unix timestamp for when the completion was created.
- `model` (string, required) — Model used for the completion.
- `choices` (list of ApiChatCompletionsPostResponsesContentApplicationJsonSchemaChoicesItems, required)
- `usage` (ApiChatCompletionsPostResponsesContentApplicationJsonSchemaUsage, optional)

## Errors

### 400 Bad Request Error

Invalid request or preprocessing failure.

- `error` (ApiChatCompletionsPostResponsesContentApplicationJsonSchemaError, required)

### 401 Unauthorized Error

No usable key for the model. The code is `activation_required` when your organization has no credits and no provider key for the model.

- `any`

### 403 Forbidden Error

The API key is missing, invalid or expired.

- `detail` (string, required)

### 404 Not Found Error

The model isn't available to your organization.

- `error` (ApiChatCompletionsPostResponsesContentApplicationJsonSchemaError, required)

### 424 Failed Dependency Error

The upstream model provider failed.

- `error` (ApiChatCompletionsPostResponsesContentApplicationJsonSchemaError, required)

## Types

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMessagesItems

- `role` (enum, required) — Message role.
  - Allowed values: `system`, `user`, `assistant`, `tool`
- `content` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMessagesItemsContent, required) — Message content. Use a string for text-only requests, or an array of content parts for multimodal requests.
- `name` (string, optional) — Optional participant name.
- `tool_call_id` (string, optional) — Required for tool response messages.

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaToolsItems

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaToolChoice

Controls tool selection. `"none"` = no tools, `"auto"` = model decides, or specify a tool object.

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLogitBias

Used to modify the probability of tokens appearing in the response

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaResponseFormat

Output format. Set `{"type": "json_schema", "json_schema": {...}}` for structured output, or `{"type": "json_object"}` for JSON mode.

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLoadBalanceGroup

Load balance group selection. Use `{"group_id": "..."}` to route through a configured group.

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCustomerCredentials

Per-customer LLM provider credentials. Keys are provider names, values are API keys.

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCredentialOverride

One-off credential overrides per provider. Overrides uploaded provider keys for this request only.

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCacheOptions

Cache behavior options. Properties: `cache_by_customer`, `is_cached_by_model`, `omit_log`.

- `cache_by_customer` (boolean, optional, default: false) — Partition cache entries by customer identifier.
- `is_cached_by_model` (boolean, optional, default: false) — Partition cache entries by model name.
- `omit_log` (boolean, optional, default: false) — Suppress log creation for cache hits.

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaPrompt

Prompt template config. Properties: `prompt_id` (required), `variables` (template variables), `version` (number, or `"latest"` for draft), `echo` (return rendered prompt), `override` (use override_params), `override_params` (OpenAI params to override), `schema_version` (`1` = legacy, `2` = prompt config wins). See [Prompt management](/docs/documentation/features/prompt-management/advanced).

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaRetryParams

Has no effect on this endpoint: chat completions always uses your organization's retry settings. Per-request `retry_params` (`retry_enabled`, `num_retries`, `retry_after`) applies only to `POST /api/responses`. See [Retries and fallback](/docs/documentation/features/gateway/retries).

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaModelNameMap

Azure deployment name mapping. Maps your custom Azure deployment names to standard model names.

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMetadata

Custom key-value metadata attached to the span.

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCustomerParams

Customer details. Properties: `customer_identifier` (takes precedence over the top-level `customer_identifier`), `name` and `email` (logged with the request, and saved on the customer when Respan first sees it), and `rate_limit` (requests per minute for this customer, overriding your organization's customer rate limit; requests over it get `429`). Budget fields sent here aren't saved or enforced. Set budgets with [Update a user](/docs/apis/users/update-user).

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLoadBalanceModelsItems

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchema

Typed metadata preserving native types (numbers, booleans, nested objects). Unlike `metadata` which coerces to strings.

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaRespanParams

Namespaced container for all Respan parameters. Alternative to passing them at top level. Must be a JSON object, or a JSON string that decodes to one. A list, number or boolean is ignored and none of its params apply.

### ApiChatCompletionsPostResponsesContentApplicationJsonSchemaChoicesItems

- `index` (integer, optional)
- `message` (ApiChatCompletionsPostResponsesContentApplicationJsonSchemaChoicesItemsMessage, optional)
- `finish_reason` (string, optional)

### ApiChatCompletionsPostResponsesContentApplicationJsonSchemaUsage

- `prompt_tokens` (integer, optional)
- `completion_tokens` (integer, optional)
- `total_tokens` (integer, optional)

### ApiChatCompletionsPostResponsesContentApplicationJsonSchemaError

- `message` (string, required)
- `type` (string, required)
- `param` (any, optional, nullable)
- `code` (any, optional, nullable)

### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMessagesItemsContent

Message content. Use a string for text-only requests, or an array of content parts for multimodal requests.

### ApiChatCompletionsPostResponsesContentApplicationJsonSchemaChoicesItemsMessage

- `role` (string, optional)
- `content` (string, optional)

## Examples

**Request**

```json
{
  "messages": [
    {
      "role": "user",
      "content": "Reply with exactly ok."
    }
  ],
  "model": "gpt-4o-mini",
  "max_tokens": 16,
  "temperature": 0
}
```

**Response**

```json
{
  "id": "chatcmpl_abc123",
  "object": "chat.completion",
  "created": 1709155200,
  "model": "gpt-4o-mini",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "ok"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 12,
    "completion_tokens": 1,
    "total_tokens": 13
  }
}
```

**SDK Code**

```python Gateway_createChatCompletion_example
import requests

url = "https://api.respan.ai/api/chat/completions"

payload = {
    "messages": [
        {
            "role": "user",
            "content": "Reply with exactly ok."
        }
    ],
    "model": "gpt-4o-mini",
    "max_tokens": 16,
    "temperature": 0
}
headers = {
    "X-Respan-Route-Provider": "vertex_ai",
    "Authorization": "Bearer <respanApiKey>",
    "Content-Type": "application/json"
}

response = requests.post(url, json=payload, headers=headers)

print(response.json())
```

```javascript Gateway_createChatCompletion_example
const url = 'https://api.respan.ai/api/chat/completions';
const options = {
  method: 'POST',
  headers: {
    'X-Respan-Route-Provider': 'vertex_ai',
    Authorization: 'Bearer <respanApiKey>',
    'Content-Type': 'application/json'
  },
  body: '{"messages":[{"role":"user","content":"Reply with exactly ok."}],"model":"gpt-4o-mini","max_tokens":16,"temperature":0}'
};

try {
  const response = await fetch(url, options);
  const data = await response.json();
  console.log(data);
} catch (error) {
  console.error(error);
}
```

```go Gateway_createChatCompletion_example
package main

import (
	"fmt"
	"strings"
	"net/http"
	"io"
)

func main() {

	url := "https://api.respan.ai/api/chat/completions"

	payload := strings.NewReader("{\n  \"messages\": [\n    {\n      \"role\": \"user\",\n      \"content\": \"Reply with exactly ok.\"\n    }\n  ],\n  \"model\": \"gpt-4o-mini\",\n  \"max_tokens\": 16,\n  \"temperature\": 0\n}")

	req, _ := http.NewRequest("POST", url, payload)

	req.Header.Add("X-Respan-Route-Provider", "vertex_ai")
	req.Header.Add("Authorization", "Bearer <respanApiKey>")
	req.Header.Add("Content-Type", "application/json")

	res, _ := http.DefaultClient.Do(req)

	defer res.Body.Close()
	body, _ := io.ReadAll(res.Body)

	fmt.Println(res)
	fmt.Println(string(body))

}
```

```ruby Gateway_createChatCompletion_example
require 'uri'
require 'net/http'

url = URI("https://api.respan.ai/api/chat/completions")

http = Net::HTTP.new(url.host, url.port)
http.use_ssl = true

request = Net::HTTP::Post.new(url)
request["X-Respan-Route-Provider"] = 'vertex_ai'
request["Authorization"] = 'Bearer <respanApiKey>'
request["Content-Type"] = 'application/json'
request.body = "{\n  \"messages\": [\n    {\n      \"role\": \"user\",\n      \"content\": \"Reply with exactly ok.\"\n    }\n  ],\n  \"model\": \"gpt-4o-mini\",\n  \"max_tokens\": 16,\n  \"temperature\": 0\n}"

response = http.request(request)
puts response.read_body
```

```java Gateway_createChatCompletion_example
import com.mashape.unirest.http.HttpResponse;
import com.mashape.unirest.http.Unirest;

HttpResponse<String> response = Unirest.post("https://api.respan.ai/api/chat/completions")
  .header("X-Respan-Route-Provider", "vertex_ai")
  .header("Authorization", "Bearer <respanApiKey>")
  .header("Content-Type", "application/json")
  .body("{\n  \"messages\": [\n    {\n      \"role\": \"user\",\n      \"content\": \"Reply with exactly ok.\"\n    }\n  ],\n  \"model\": \"gpt-4o-mini\",\n  \"max_tokens\": 16,\n  \"temperature\": 0\n}")
  .asString();
```

```php Gateway_createChatCompletion_example
<?php
require_once('vendor/autoload.php');

$client = new \GuzzleHttp\Client();

$response = $client->request('POST', 'https://api.respan.ai/api/chat/completions', [
  'body' => '{
  "messages": [
    {
      "role": "user",
      "content": "Reply with exactly ok."
    }
  ],
  "model": "gpt-4o-mini",
  "max_tokens": 16,
  "temperature": 0
}',
  'headers' => [
    'Authorization' => 'Bearer <respanApiKey>',
    'Content-Type' => 'application/json',
    'X-Respan-Route-Provider' => 'vertex_ai',
  ],
]);

echo $response->getBody();
```

```csharp Gateway_createChatCompletion_example
using RestSharp;

var client = new RestClient("https://api.respan.ai/api/chat/completions");
var request = new RestRequest(Method.POST);
request.AddHeader("X-Respan-Route-Provider", "vertex_ai");
request.AddHeader("Authorization", "Bearer <respanApiKey>");
request.AddHeader("Content-Type", "application/json");
request.AddParameter("application/json", "{\n  \"messages\": [\n    {\n      \"role\": \"user\",\n      \"content\": \"Reply with exactly ok.\"\n    }\n  ],\n  \"model\": \"gpt-4o-mini\",\n  \"max_tokens\": 16,\n  \"temperature\": 0\n}", ParameterType.RequestBody);
IRestResponse response = client.Execute(request);
```

```swift Gateway_createChatCompletion_example
import Foundation

let headers = [
  "X-Respan-Route-Provider": "vertex_ai",
  "Authorization": "Bearer <respanApiKey>",
  "Content-Type": "application/json"
]
let parameters = [
  "messages": [
    [
      "role": "user",
      "content": "Reply with exactly ok."
    ]
  ],
  "model": "gpt-4o-mini",
  "max_tokens": 16,
  "temperature": 0
] as [String : Any]

let postData = JSONSerialization.data(withJSONObject: parameters, options: [])

let request = NSMutableURLRequest(url: NSURL(string: "https://api.respan.ai/api/chat/completions")! as URL,
                                        cachePolicy: .useProtocolCachePolicy,
                                    timeoutInterval: 10.0)
request.httpMethod = "POST"
request.allHTTPHeaderFields = headers
request.httpBody = postData as Data

let session = URLSession.shared
let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in
  if (error != nil) {
    print(error as Any)
  } else {
    let httpResponse = response as? HTTPURLResponse
    print(httpResponse)
  }
})

dataTask.resume()
```