> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://www.respan.ai/docs/apis/gateway/create-chat-completion/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://www.respan.ai/_mcp/server. # Create a chat completion POST https://api.respan.ai/api/chat/completions Content-Type: application/json Sends a chat completion request through the Respan gateway with automatic logging. Accepts [OpenAI chat completion parameters](https://platform.openai.com/docs/apis/chat) and Respan options for fallbacks, caching, and prompt management. Pass Respan parameters in top-level body fields, under `respan_params`, or as base64-encoded JSON in the `X-Data-Respan-Params` header. If the same field is sent more than one way, `respan_params` takes precedence over the header, and both take precedence over a top-level field. The exception is `variables`: top-level `variables` are merged key by key over the `variables` from `respan_params` or the header. `respan_params`, `keywordsai_params`, and the decoded header must each be a JSON object; a top-level `respan_params` may also be a JSON string that decodes to one. Any other value is ignored, and so are the header's parameters: the request is served, but none of those parameters apply. Fields sent at the top level still apply. With the OpenAI SDK, use `extra_body`. For legacy compatibility, `keywordsai_params` is merged into `respan_params`, and `X-Data-Keywordsai-Params` is still accepted and renamed internally. Reference: https://www.respan.ai/docs/apis/gateway/create-chat-completion ## Authentication - `Authorization` header (bearer token, required) — Use your Respan API key for Respan API authentication. Enter only the Respan API key value; clients send Authorization: Bearer \. For /api/responses, provider credentials such as Perplexity, OpenAI, or Azure OpenAI go in Settings -> Providers or respan\_params.credential\_override in the request body, not in this authentication field. ## Request ### Headers - `X-Data-Respan-Params` (string, optional) — Base64-encoded JSON object of Respan parameters. Legacy `X-Data-Keywordsai-Params` is still accepted. - `X-Respan-Route-Provider` (string, optional) — Pin the request to a specific provider without changing the model slug. Example: `vertex_ai` routes a `claude-sonnet-4-5-20250929` request to Vertex AI Claude. - `X-Respan-Beta` (string, optional) — Comma-separated beta feature flags. Available: token-breakdown-2026-03-26, env-scoped-integrations-2026-03-28 ### Body (application/json) This endpoint expects an object. - `messages` (list of ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMessagesItems, required) — Array of messages in the conversation. Each message has `role` (`system`, `user`, `assistant`, `tool`) and `content`. - `model` (string, required) — Model to use. See [Models](https://platform.respan.ai/platform/models) for available options. - `stream` (boolean, optional) — Stream back partial progress token by token as server-sent events. - `tools` (list of ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaToolsItems, optional) — Tools the model may call. Currently only functions are supported. - `tool_choice` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaToolChoice, optional) — Controls tool selection. `"none"` = no tools, `"auto"` = model decides, or specify a tool object. - `frequency_penalty` (double, optional) — Penalizes tokens based on frequency in text so far (-2 to 2). - `max_tokens` (double, optional) — Maximum tokens to generate. - `temperature` (double, optional, default: 1) — Sampling temperature (0-2). Higher = more random. - `n` (double, optional, default: 1) — Number of completions to generate. Note: costs multiply with `n`. - `logprobs` (boolean, optional) — Return log probabilities of output tokens. - `echo` (boolean, optional) — Echo back the prompt in addition to the completion - `stop` (list of string, optional) — Stop sequences where generation halts. - `presence_penalty` (double, optional) — Penalizes tokens already present in text (-2 to 2). - `logit_bias` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLogitBias, optional) — Used to modify the probability of tokens appearing in the response - `response_format` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaResponseFormat, optional) — Output format. Set `{"type": "json_schema", "json_schema": {...}}` for structured output, or `{"type": "json_object"}` for JSON mode. - `parallel_tool_calls` (boolean, optional) — Enable parallel function calling during tool use. - `load_balance_group` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLoadBalanceGroup, optional) — Load balance group selection. Use `{"group_id": "..."}` to route through a configured group. - `fallback_models` (list of string, optional) — Backup models (ranked by priority) if the primary model fails. - `customer_credentials` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCustomerCredentials, optional) — Per-customer LLM provider credentials. Keys are provider names, values are API keys. - `credential_override` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCredentialOverride, optional) — One-off credential overrides per provider. Overrides uploaded provider keys for this request only. - `cache_enabled` (boolean, optional) — Enable response caching. See [Caching](/docs/documentation/features/gateway/caching). - `cache_ttl` (double, optional) — How long a cached response is served, in seconds. A response stored without `cache_ttl` is served from the cache for at most 30 minutes. - `cache_options` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCacheOptions, optional) — Cache behavior options. Properties: `cache_by_customer`, `is_cached_by_model`, `omit_log`. - `prompt` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaPrompt, optional) — Prompt template config. Properties: `prompt_id` (required), `variables` (template variables), `version` (number, or `"latest"` for draft), `echo` (return rendered prompt), `override` (use override_params), `override_params` (OpenAI params to override), `schema_version` (`1` = legacy, `2` = prompt config wins). See [Prompt management](/docs/documentation/features/prompt-management/advanced). - `retry_params` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaRetryParams, optional) — Has no effect on this endpoint: chat completions always uses your organization's retry settings. Per-request `retry_params` (`retry_enabled`, `num_retries`, `retry_after`) applies only to `POST /api/responses`. See [Retries and fallback](/docs/documentation/features/gateway/retries). - `disable_log` (boolean, optional) — When `true`, omits input/output from the log. Metrics (tokens, cost, latency) are still recorded. - `model_name_map` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaModelNameMap, optional) — Azure deployment name mapping. Maps your custom Azure deployment names to standard model names. - `models` (list of string, optional) — Model list for LLM router selection. - `exclude_providers` (list of string, optional) — Providers to exclude from routing. All models under excluded providers are skipped. - `exclude_models` (list of string, optional) — Specific models to exclude from routing. - `metadata` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMetadata, optional) — Custom key-value metadata attached to the span. - `custom_identifier` (string, optional) — Indexed custom tag for fast querying. - `customer_identifier` (string, optional) — End user identifier for analytics and budgets. - `customer_params` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCustomerParams, optional) — Customer details. Properties: `customer_identifier` (takes precedence over the top-level `customer_identifier`), `name` and `email` (logged with the request, and saved on the customer when Respan first sees it), and `rate_limit` (requests per minute for this customer, overriding your organization's customer rate limit; requests over it get `429`). Budget fields sent here aren't saved or enforced. Set budgets with [Update a user](/docs/apis/users/update-user). - `request_breakdown` (boolean, optional) — Return response metrics summary in the response body. For streaming, metrics appear in the final chunk. - `positive_feedback` (boolean, optional) — User feedback. `true` = liked, `false` = disliked. - `load_balance_models` (list of ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLoadBalanceModelsItems, optional) — Inline load balancing options. Each item can include `model`, `weight`, and optional `credentials`. - `thread_identifier` (string, optional) — Conversation thread ID. Spans with the same `thread_identifier` are grouped together. - `properties` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchema, optional) — Typed metadata preserving native types (numbers, booleans, nested objects). Unlike `metadata` which coerces to strings. - `retries` (integer, optional, default: 0) — Has no effect: chat completions always uses your organization's retry settings. See [Retries and fallback](/docs/documentation/features/gateway/retries). - `weight` (double, optional) — Load balancing weight. - `span_name` (string, optional) — Custom span name for tracing. - `respan_params` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaRespanParams, optional) — Namespaced container for all Respan parameters. Alternative to passing them at top level. Must be a JSON object, or a JSON string that decodes to one. A list, number or boolean is ignored and none of its params apply. ## Response ### 200 Successful response for Create chat completion - `id` (string, required) — Chat completion ID. - `object` (string, required) - `created` (integer, required) — Unix timestamp for when the completion was created. - `model` (string, required) — Model used for the completion. - `choices` (list of ApiChatCompletionsPostResponsesContentApplicationJsonSchemaChoicesItems, required) - `usage` (ApiChatCompletionsPostResponsesContentApplicationJsonSchemaUsage, optional) ## Errors ### 400 Bad Request Error Invalid request or preprocessing failure. - `error` (ApiChatCompletionsPostResponsesContentApplicationJsonSchemaError, required) ### 401 Unauthorized Error No usable key for the model. The code is `activation_required` when your organization has no credits and no provider key for the model. - `any` ### 403 Forbidden Error The API key is missing, invalid or expired. - `detail` (string, required) ### 404 Not Found Error The model isn't available to your organization. - `error` (ApiChatCompletionsPostResponsesContentApplicationJsonSchemaError, required) ### 424 Failed Dependency Error The upstream model provider failed. - `error` (ApiChatCompletionsPostResponsesContentApplicationJsonSchemaError, required) ## Types ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMessagesItems - `role` (enum, required) — Message role. - Allowed values: `system`, `user`, `assistant`, `tool` - `content` (ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMessagesItemsContent, required) — Message content. Use a string for text-only requests, or an array of content parts for multimodal requests. - `name` (string, optional) — Optional participant name. - `tool_call_id` (string, optional) — Required for tool response messages. ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaToolsItems ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaToolChoice Controls tool selection. `"none"` = no tools, `"auto"` = model decides, or specify a tool object. ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLogitBias Used to modify the probability of tokens appearing in the response ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaResponseFormat Output format. Set `{"type": "json_schema", "json_schema": {...}}` for structured output, or `{"type": "json_object"}` for JSON mode. ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLoadBalanceGroup Load balance group selection. Use `{"group_id": "..."}` to route through a configured group. ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCustomerCredentials Per-customer LLM provider credentials. Keys are provider names, values are API keys. ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCredentialOverride One-off credential overrides per provider. Overrides uploaded provider keys for this request only. ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCacheOptions Cache behavior options. Properties: `cache_by_customer`, `is_cached_by_model`, `omit_log`. - `cache_by_customer` (boolean, optional, default: false) — Partition cache entries by customer identifier. - `is_cached_by_model` (boolean, optional, default: false) — Partition cache entries by model name. - `omit_log` (boolean, optional, default: false) — Suppress log creation for cache hits. ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaPrompt Prompt template config. Properties: `prompt_id` (required), `variables` (template variables), `version` (number, or `"latest"` for draft), `echo` (return rendered prompt), `override` (use override_params), `override_params` (OpenAI params to override), `schema_version` (`1` = legacy, `2` = prompt config wins). See [Prompt management](/docs/documentation/features/prompt-management/advanced). ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaRetryParams Has no effect on this endpoint: chat completions always uses your organization's retry settings. Per-request `retry_params` (`retry_enabled`, `num_retries`, `retry_after`) applies only to `POST /api/responses`. See [Retries and fallback](/docs/documentation/features/gateway/retries). ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaModelNameMap Azure deployment name mapping. Maps your custom Azure deployment names to standard model names. ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMetadata Custom key-value metadata attached to the span. ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaCustomerParams Customer details. Properties: `customer_identifier` (takes precedence over the top-level `customer_identifier`), `name` and `email` (logged with the request, and saved on the customer when Respan first sees it), and `rate_limit` (requests per minute for this customer, overriding your organization's customer rate limit; requests over it get `429`). Budget fields sent here aren't saved or enforced. Set budgets with [Update a user](/docs/apis/users/update-user). ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaLoadBalanceModelsItems ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchema Typed metadata preserving native types (numbers, booleans, nested objects). Unlike `metadata` which coerces to strings. ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaRespanParams Namespaced container for all Respan parameters. Alternative to passing them at top level. Must be a JSON object, or a JSON string that decodes to one. A list, number or boolean is ignored and none of its params apply. ### ApiChatCompletionsPostResponsesContentApplicationJsonSchemaChoicesItems - `index` (integer, optional) - `message` (ApiChatCompletionsPostResponsesContentApplicationJsonSchemaChoicesItemsMessage, optional) - `finish_reason` (string, optional) ### ApiChatCompletionsPostResponsesContentApplicationJsonSchemaUsage - `prompt_tokens` (integer, optional) - `completion_tokens` (integer, optional) - `total_tokens` (integer, optional) ### ApiChatCompletionsPostResponsesContentApplicationJsonSchemaError - `message` (string, required) - `type` (string, required) - `param` (any, optional, nullable) - `code` (any, optional, nullable) ### ApiChatCompletionsPostRequestBodyContentApplicationJsonSchemaMessagesItemsContent Message content. Use a string for text-only requests, or an array of content parts for multimodal requests. ### ApiChatCompletionsPostResponsesContentApplicationJsonSchemaChoicesItemsMessage - `role` (string, optional) - `content` (string, optional) ## Examples **Request** ```json { "messages": [ { "role": "user", "content": "Reply with exactly ok." } ], "model": "gpt-4o-mini", "max_tokens": 16, "temperature": 0 } ``` **Response** ```json { "id": "chatcmpl_abc123", "object": "chat.completion", "created": 1709155200, "model": "gpt-4o-mini", "choices": [ { "index": 0, "message": { "role": "assistant", "content": "ok" }, "finish_reason": "stop" } ], "usage": { "prompt_tokens": 12, "completion_tokens": 1, "total_tokens": 13 } } ``` **SDK Code** ```python Gateway_createChatCompletion_example import requests url = "https://api.respan.ai/api/chat/completions" payload = { "messages": [ { "role": "user", "content": "Reply with exactly ok." } ], "model": "gpt-4o-mini", "max_tokens": 16, "temperature": 0 } headers = { "X-Respan-Route-Provider": "vertex_ai", "Authorization": "Bearer ", "Content-Type": "application/json" } response = requests.post(url, json=payload, headers=headers) print(response.json()) ``` ```javascript Gateway_createChatCompletion_example const url = 'https://api.respan.ai/api/chat/completions'; const options = { method: 'POST', headers: { 'X-Respan-Route-Provider': 'vertex_ai', Authorization: 'Bearer ', 'Content-Type': 'application/json' }, body: '{"messages":[{"role":"user","content":"Reply with exactly ok."}],"model":"gpt-4o-mini","max_tokens":16,"temperature":0}' }; try { const response = await fetch(url, options); const data = await response.json(); console.log(data); } catch (error) { console.error(error); } ``` ```go Gateway_createChatCompletion_example package main import ( "fmt" "strings" "net/http" "io" ) func main() { url := "https://api.respan.ai/api/chat/completions" payload := strings.NewReader("{\n \"messages\": [\n {\n \"role\": \"user\",\n \"content\": \"Reply with exactly ok.\"\n }\n ],\n \"model\": \"gpt-4o-mini\",\n \"max_tokens\": 16,\n \"temperature\": 0\n}") req, _ := http.NewRequest("POST", url, payload) req.Header.Add("X-Respan-Route-Provider", "vertex_ai") req.Header.Add("Authorization", "Bearer ") req.Header.Add("Content-Type", "application/json") res, _ := http.DefaultClient.Do(req) defer res.Body.Close() body, _ := io.ReadAll(res.Body) fmt.Println(res) fmt.Println(string(body)) } ``` ```ruby Gateway_createChatCompletion_example require 'uri' require 'net/http' url = URI("https://api.respan.ai/api/chat/completions") http = Net::HTTP.new(url.host, url.port) http.use_ssl = true request = Net::HTTP::Post.new(url) request["X-Respan-Route-Provider"] = 'vertex_ai' request["Authorization"] = 'Bearer ' request["Content-Type"] = 'application/json' request.body = "{\n \"messages\": [\n {\n \"role\": \"user\",\n \"content\": \"Reply with exactly ok.\"\n }\n ],\n \"model\": \"gpt-4o-mini\",\n \"max_tokens\": 16,\n \"temperature\": 0\n}" response = http.request(request) puts response.read_body ``` ```java Gateway_createChatCompletion_example import com.mashape.unirest.http.HttpResponse; import com.mashape.unirest.http.Unirest; HttpResponse response = Unirest.post("https://api.respan.ai/api/chat/completions") .header("X-Respan-Route-Provider", "vertex_ai") .header("Authorization", "Bearer ") .header("Content-Type", "application/json") .body("{\n \"messages\": [\n {\n \"role\": \"user\",\n \"content\": \"Reply with exactly ok.\"\n }\n ],\n \"model\": \"gpt-4o-mini\",\n \"max_tokens\": 16,\n \"temperature\": 0\n}") .asString(); ``` ```php Gateway_createChatCompletion_example request('POST', 'https://api.respan.ai/api/chat/completions', [ 'body' => '{ "messages": [ { "role": "user", "content": "Reply with exactly ok." } ], "model": "gpt-4o-mini", "max_tokens": 16, "temperature": 0 }', 'headers' => [ 'Authorization' => 'Bearer ', 'Content-Type' => 'application/json', 'X-Respan-Route-Provider' => 'vertex_ai', ], ]); echo $response->getBody(); ``` ```csharp Gateway_createChatCompletion_example using RestSharp; var client = new RestClient("https://api.respan.ai/api/chat/completions"); var request = new RestRequest(Method.POST); request.AddHeader("X-Respan-Route-Provider", "vertex_ai"); request.AddHeader("Authorization", "Bearer "); request.AddHeader("Content-Type", "application/json"); request.AddParameter("application/json", "{\n \"messages\": [\n {\n \"role\": \"user\",\n \"content\": \"Reply with exactly ok.\"\n }\n ],\n \"model\": \"gpt-4o-mini\",\n \"max_tokens\": 16,\n \"temperature\": 0\n}", ParameterType.RequestBody); IRestResponse response = client.Execute(request); ``` ```swift Gateway_createChatCompletion_example import Foundation let headers = [ "X-Respan-Route-Provider": "vertex_ai", "Authorization": "Bearer ", "Content-Type": "application/json" ] let parameters = [ "messages": [ [ "role": "user", "content": "Reply with exactly ok." ] ], "model": "gpt-4o-mini", "max_tokens": 16, "temperature": 0 ] as [String : Any] let postData = JSONSerialization.data(withJSONObject: parameters, options: []) let request = NSMutableURLRequest(url: NSURL(string: "https://api.respan.ai/api/chat/completions")! as URL, cachePolicy: .useProtocolCachePolicy, timeoutInterval: 10.0) request.httpMethod = "POST" request.allHTTPHeaderFields = headers request.httpBody = postData as Data let session = URLSession.shared let dataTask = session.dataTask(with: request as URLRequest, completionHandler: { (data, response, error) -> Void in if (error != nil) { print(error as Any) } else { let httpResponse = response as? HTTPURLResponse print(httpResponse) } }) dataTask.resume() ```