Free
$0
- Bionic agent, local models on llama.cpp and MLX, offline voice transcription, LM Link for up to 5 devices, limited web search
LM Studio is a desktop app from Element Labs for downloading and running open models on your own computer, with llama.cpp and Apple's MLX as its engines. Its newest release centers on Bionic, an agent that creates documents, slides, PDFs, and code using local models or open models hosted in LM Studio's own US cloud.
For developers, LM Studio runs a local server with OpenAI- and Anthropic-compatible endpoints alongside its own REST API, plus TypeScript and Python SDKs and the lms CLI. A headless daemon called llmster runs the same core on servers or in CI without the GUI, and LM Link lets one device use models loaded on another, for up to five devices on the free plan.
Running models locally is free, and the paid plans add usage of US-hosted open models with zero data retention. Team subscriptions aren't available yet, only centralized billing for credits. LM Studio is a runtime and agent rather than a production platform, so routing across providers, evals, and production tracing come from other tools.
Core capabilities this platform advertises.
What this tool does well, and the limitations to keep in mind.
Pros
Cons
What's included in each plan, and how the tiers compare.
$0
$20
Per month
$100
Per month
Developers and individuals who want to run open models on their own computer through a desktop app or a local API
LM Studio runs open models on your machine. When an app built on those models moves to production, Respan routes across 1,000+ hosted models, including open models from providers like Together AI, Fireworks, and Novita, with tracing, spend limits, and evals on every call.
Top companies in Inference & Compute you can use instead of LM Studio.
Side-by-side comparisons with other tools in this category.
Companies from adjacent layers in the AI stack that work well with LM Studio.
Respan routes across 1,000+ hosted models, including open models from providers like Together AI, Fireworks, and Novita, with tracing, spend limits, and evals on live traffic built in. Start free.
llama.cpp
GGUF universal model format (weights + tokenizer + metadata in one file)
llama.cpp vs LM Studio