Free
$0
- Unlimited local use, starter usage credits for starter cloud models, 1 concurrent cloud request
Ollama is an open-source tool from Ollama Inc. for running open models like Gemma, Qwen models, DeepSeek, and gpt-oss on your own machine through a CLI and a local REST API. It's MIT-licensed, installs on macOS, Windows, and Linux or runs in Docker, and uses llama.cpp as a supported backend.
Ollama also runs models in its own cloud. Free accounts get starter usage credits, paid plans add monthly credits and higher concurrency, and running models on your own hardware stays unlimited on every plan. Ollama states that prompts and responses are never logged or trained on, and the ollama launch command connects models directly to coding agents like Codex, OpenCode, and Claude Code, which makes it easy to run open models on either side of an OpenCode vs Claude Code comparison.
Ollama is built to run models rather than operate them in production. It doesn't route across providers or score outputs, and observability comes from third-party tools that integrate with it.
Core capabilities this platform advertises.
What this tool does well, and the limitations to keep in mind.
Pros
Cons
What's included in each plan, and how the tiers compare.
$0
$20/mo or $200/yr
Per month or per year
$100
Per month
Developers who want to run open models locally for prototyping, private workloads, or local coding agents
Respan has a native Ollama integration. The respan-instrumentation-ollama package patches the official Ollama Python client and sends chat, generation, tool call, and embedding spans to Respan, so you keep running models locally and get tracing, evals, and cost tracking on every call.
Top companies in Inference & Compute you can use instead of Ollama.
Side-by-side comparisons with other tools in this category.
Companies from adjacent layers in the AI stack that work well with Ollama.
$500
Per month
Custom
Contact sales for a quote
Respan's Ollama integration traces every chat, generation, and embedding call from the official Python client, so you can score local model output with real evals. When you're ready for hosted models, route across 1,000+ of them from the same account. Start free.
llama.cpp
GGUF universal model format (weights + tokenizer + metadata in one file)
llama.cpp vs Ollama