Create an OpenAI-format chat completion
Relay an OpenAI-format chat completions payload to one provider and get that provider’s response back byte for byte. Respan authenticates the request, applies your limits, and logs it, but doesn’t translate the request or recompute usage. Respan removes respan_params and merges a literal extra_body object into the body; unknown and provider-only fields are passed through. Streams are relayed frame by frame, and requests time out after 600 seconds.
{provider} is a provider with an OpenAI-format API, such as openai, azure_openai, xai, groq, mistral, deepseek, togetherai, fireworks, or perplexity. A provider without an OpenAI-format API returns 404. A provider with no fixed public endpoint, such as Azure OpenAI, needs the provider URL saved with its key; without it the request returns 400. Anthropic, Google and OpenRouter have their own endpoints.
Upstream, Respan sends only accept, accept-language, content-type, user-agent, and headers that start with x- plus the provider ID, with _ written as - (for example x-azure-openai-). The response includes X-Respan-Log-Id, X-Respan-Upstream-Request-Id, and the provider’s x-ratelimit-* and retry-after headers.
Authentication
Use your Respan API key for Respan API authentication. Enter only the Respan API key value; clients send Authorization: Bearer <RESPAN_API_KEY>. For /api/responses, provider credentials such as Perplexity, OpenAI, or Azure OpenAI go in Settings -> Providers or respan_params.credential_override in the request body, not in this authentication field.
Path parameters
The provider ID, such as openai, azure_openai, xai, groq, mistral, deepseek, togetherai, fireworks, or perplexity.
Request
Stream the provider's response as server-sent events.
Respan parameters such as customer_identifier or metadata. Removed before the request is forwarded.