Gemma 4 31b is available on the Respan AI Gateway for production LLM workloads. Supports explicit caching and implicit caching. Up to 33K context window.
from openai import OpenAI client = OpenAI( base_url="https://api.respan.ai/api/", api_key="YOUR_RESPAN_API_KEY",) response = client.chat.completions.create( model="prism/gemma-4-31b", messages=[{"role": "user", "content": "Hello!"}],)print(response.choices[0].message.content)from openai import OpenAI client = OpenAI( base_url="https://api.respan.ai/api/", api_key="YOUR_RESPAN_API_KEY",) response = client.chat.completions.create( model="prism/gemma-4-31b", messages=[{"role": "user", "content": "Hello!"}],)print(response.choices[0].message.content)Gateway routes for this model on the Respan gateway, including fallback routes when configured.
| Model | Capabilities | ||||||
|---|---|---|---|---|---|---|---|
| prism/gemma-4-31b | 33K | $0.14/M | $0.40/M | Read:$0.06/MWrite:$0.00/M | — | — |
Other models from the same provider available through the gateway.
| Model | Capabilities | ||||||
|---|---|---|---|---|---|---|---|
| prism/deepseek-v4-flashSave 20% | 1M | $0.02/M | $0.44/M | Read:$0.01/MWrite:$0.00/M | 7.3s | 88tps | |
| prism/deepseek-v4.1-flashSave 20% | 1M | $0.14/M | $0.50/M | Read:$0.0048/MWrite:$0.00/M | 3.0s | +2 | 244tps |
Built to meet the security and privacy standards that enterprise and healthcare teams require.
ISO 27001
The internationally recognized standard for information security management.
SOC 2
Secure, compliant management of your data across all of our systems.
GDPR
Operated under GDPR, the world's strictest standard for data privacy.
HIPAA
HIPAA compliant, with a BAA available for healthcare teams.