Skip to navigation

Create a chat completion

Sends a chat completion request through the Respan gateway with automatic logging. Accepts OpenAI chat completion parameters and Respan options for fallbacks, caching, and prompt management.

Pass Respan parameters in top-level body fields, under respan_params, or as base64-encoded JSON in the X-Data-Respan-Params header. If the same field is sent more than one way, respan_params takes precedence over the header, and both take precedence over a top-level field. The exception is variables: top-level variables are merged key by key over the variables from respan_params or the header. respan_params, keywordsai_params, and the decoded header must each be a JSON object; a top-level respan_params may also be a JSON string that decodes to one. Any other value is ignored, and so are the header’s parameters: the request is served, but none of those parameters apply. Fields sent at the top level still apply. With the OpenAI SDK, use extra_body.

For legacy compatibility, keywordsai_params is merged into respan_params, and X-Data-Keywordsai-Params is still accepted and renamed internally.

Authentication

AuthorizationBearer

Use your Respan API key for Respan API authentication. Enter only the Respan API key value; clients send Authorization: Bearer <RESPAN_API_KEY>. For /api/responses, provider credentials such as Perplexity, OpenAI, or Azure OpenAI go in Settings -> Providers or respan_params.credential_override in the request body, not in this authentication field.

Headers

X-Data-Respan-ParamsstringOptional

Base64-encoded JSON object of Respan parameters. Legacy X-Data-Keywordsai-Params is still accepted.

X-Respan-Route-ProviderstringOptional

Pin the request to a specific provider without changing the model slug. Example: vertex_ai routes a claude-sonnet-4-5-20250929 request to Vertex AI Claude.

X-Respan-BetastringOptional

Comma-separated beta feature flags. Available: token-breakdown-2026-03-26, env-scoped-integrations-2026-03-28

Request

This endpoint expects an object.
messageslist of objectsRequired

Array of messages in the conversation. Each message has role (system, user, assistant, tool) and content.

modelstringRequired

Model to use. See Models for available options.

streambooleanOptional

Stream back partial progress token by token as server-sent events.

toolslist of objectsOptional
Tools the model may call. Currently only functions are supported.
tool_choiceobjectOptional

Controls tool selection. "none" = no tools, "auto" = model decides, or specify a tool object.

frequency_penaltydoubleOptional

Penalizes tokens based on frequency in text so far (-2 to 2).

max_tokensdoubleOptional
Maximum tokens to generate.
temperaturedoubleOptionalDefaults to 1

Sampling temperature (0-2). Higher = more random.

ndoubleOptionalDefaults to 1

Number of completions to generate. Note: costs multiply with n.

logprobsbooleanOptional
Return log probabilities of output tokens.
echobooleanOptional
Echo back the prompt in addition to the completion
stoplist of stringsOptional
Stop sequences where generation halts.
presence_penaltydoubleOptional

Penalizes tokens already present in text (-2 to 2).

logit_biasobjectOptional
Used to modify the probability of tokens appearing in the response
response_formatobjectOptional

Output format. Set {"type": "json_schema", "json_schema": {...}} for structured output, or {"type": "json_object"} for JSON mode.

parallel_tool_callsbooleanOptional
Enable parallel function calling during tool use.
load_balance_groupobjectOptional

Load balance group selection. Use {"group_id": "..."} to route through a configured group.

fallback_modelslist of stringsOptional

Backup models (ranked by priority) if the primary model fails.

customer_credentialsobjectOptional

Per-customer LLM provider credentials. Keys are provider names, values are API keys.

credential_overrideobjectOptional

One-off credential overrides per provider. Overrides uploaded provider keys for this request only.

cache_enabledbooleanOptional

Enable response caching. See Caching.

cache_ttldoubleOptional

How long a cached response is served, in seconds. A response stored without cache_ttl is served from the cache for at most 30 minutes.

cache_optionsobjectOptional

Cache behavior options. Properties: cache_by_customer, is_cached_by_model, omit_log.

promptobjectOptional

Prompt template config. Properties: prompt_id (required), variables (template variables), version (number, or "latest" for draft), echo (return rendered prompt), override (use override_params), override_params (OpenAI params to override), schema_version (1 = legacy, 2 = prompt config wins). See Prompt management.

retry_paramsobjectOptional

Has no effect on this endpoint: chat completions always uses your organization's retry settings. Per-request retry_params (retry_enabled, num_retries, retry_after) applies only to POST /api/responses. See Retries and fallback.

disable_logbooleanOptional

When true, omits input/output from the log. Metrics (tokens, cost, latency) are still recorded.

model_name_mapobjectOptional
Azure deployment name mapping. Maps your custom Azure deployment names to standard model names.
modelslist of stringsOptional
Model list for LLM router selection.
exclude_providerslist of stringsOptional
Providers to exclude from routing. All models under excluded providers are skipped.
exclude_modelslist of stringsOptional
Specific models to exclude from routing.
metadataobjectOptional

Custom key-value metadata attached to the span.

custom_identifierstringOptional
Indexed custom tag for fast querying.
customer_identifierstringOptional<=254 characters
End user identifier for analytics and budgets.
customer_paramsobjectOptional

Customer details. Properties: customer_identifier (takes precedence over the top-level customer_identifier), name and email (logged with the request, and saved on the customer when Respan first sees it), and rate_limit (requests per minute for this customer, overriding your organization's customer rate limit; requests over it get 429). Budget fields sent here aren't saved or enforced. Set budgets with Update a user.

request_breakdownbooleanOptional
Return response metrics summary in the response body. For streaming, metrics appear in the final chunk.
positive_feedbackbooleanOptional

User feedback. true = liked, false = disliked.

load_balance_modelslist of objectsOptional

Inline load balancing options. Each item can include model, weight, and optional credentials.

thread_identifierstringOptional

Conversation thread ID. Spans with the same thread_identifier are grouped together.

propertiesobjectOptional

Typed metadata preserving native types (numbers, booleans, nested objects). Unlike metadata which coerces to strings.

retriesintegerOptionalDefaults to 0

Has no effect: chat completions always uses your organization's retry settings. See Retries and fallback.

weightdoubleOptional
Load balancing weight.
span_namestringOptional
Custom span name for tracing.
respan_paramsobjectOptional
Namespaced container for all Respan parameters. Alternative to passing them at top level. Must be a JSON object, or a JSON string that decodes to one. A list, number or boolean is ignored and none of its params apply.

Response

Successful response for Create chat completion
idstring
Chat completion ID.
objectstring
createdinteger
Unix timestamp for when the completion was created.
modelstring
Model used for the completion.
choiceslist of objects
usageobjectOptional

Errors

400
Bad Request Error
401
Unauthorized Error
403
Forbidden Error
404
Not Found Error
424
Failed Dependency Error