Ollama
Runs open-weight models on your own machine with one command, and optionally spills over to Ollama's cloud when the model is too big for local hardware.
Updated 2026-08-02
Yes, within limits. Ollama runs a free tier with paid plans above it.
Listed pricing: Free, with Pro at $20/mo and Team at $25/seat/mo for cloud usage.
Overview
Ollama calls itself "The easiest way to build with open models", and the front page frames the job as getting "up and running with OpenClaw, Claude Code, and more in minutes using open models". That is the accurate description of what it has become. It is less a chat app than a runtime, the thing other tools point at when they want an open-weight model served locally. Downloads exist for macOS, which requires macOS 14 Sonoma or later, plus Windows and Linux.
The local-first framing is the whole product. Models run on your hardware, work continues offline, and nothing is sent anywhere unless you pick a cloud model. Ollama has since added that cloud tier for models too large to fit on a laptop, served from US, European and Singaporean regions, with the stated policy that "Prompt or response data is never logged or trained on". The free account covers local inference, unlimited public models and what Ollama describes as light cloud usage with one concurrent model.
My read: the pricing is honest about the model but vague about the quantity. Pro at $20 a month buys "50x more cloud usage than Free" and three concurrent cloud models, Max at $100 a month buys ten concurrent models and "5x more usage than Pro", and the $100 Max tier is currently marked as paused for new sign-ups. Every one of those numbers is a multiplier against a base that Ollama never publishes, so the only way to size a plan is to run the workload and watch. For local-only use it does not matter, because local inference costs nothing but electricity and RAM. It matters the moment you rely on the cloud path.
Is Ollama free?
Yes, within limits. Ollama runs a free tier with paid plans above it.
What the free tier covers: Yes, and it is the main event. Local inference is free forever on your own hardware. The paid tiers exist to buy cloud capacity, and their allowances are published only as multipliers of an unstated base.
Listed pricing: Free, with Pro at $20/mo and Team at $25/seat/mo for cloud usage.
Pricing verified . That is the date the plans were last re-checked against the vendor's own pages, not today's date. Confirm on the official site before you pay.
What the free tier leaves out
Read straight off the plan list below. Vendors move features between tiers, so check the current split before you pay.
- Pro $20/mo or $200/yr Larger cloud models, three cloud models at a time, "50x more cloud usage than Free", and the ability to upload and share private models.
- Max $100/mo Ten cloud models at a time and "5x more usage than Pro". Listed as paused for new sign-ups.
- Team $25/seat/mo, five-seat minimum Open models served in the US and Europe, zero data retention and logging, shared billing and administration.
- Enterprise Custom Custom terms for larger organisations.
Ollama pricing
| Plan | Price | What's included |
|---|---|---|
| Free | $0 | Local inference on your own hardware, unlimited public models, and light cloud usage with one concurrent cloud model. |
| Pro | $20/mo or $200/yr | Larger cloud models, three cloud models at a time, "50x more cloud usage than Free", and the ability to upload and share private models. |
| Max | $100/mo | Ten cloud models at a time and "5x more usage than Pro". Listed as paused for new sign-ups. |
| Team | $25/seat/mo, five-seat minimum | Open models served in the US and Europe, zero data retention and logging, shared billing and administration. |
| Enterprise | Custom | Custom terms for larger organisations. |
Local inference on your own hardware, unlimited public models, and light cloud usage with one concurrent cloud model.
Larger cloud models, three cloud models at a time, "50x more cloud usage than Free", and the ability to upload and share private models.
Ten cloud models at a time and "5x more usage than Pro". Listed as paused for new sign-ups.
Open models served in the US and Europe, zero data retention and logging, shared billing and administration.
Custom terms for larger organisations.
One plan carries no published number: Enterprise, recorded as “Custom”. Nothing is estimated in its place. The vendor's own pricing page is the only source for those figures.
Pricing verified . That is the date the plans were last re-checked against the vendor's own pages, not today's date. Confirm on the official site before you pay.
Is Ollama worth it?
Worth it for Anyone who wants open-weight models running on their own hardware for privacy, offline work or zero marginal cost, with a cloud fallback for models that will not fit.
You can test that on the free tier before paying anything. The recorded trade-offs are listed below, and any one of them can settle the question on its own.
The 8.8/10 AI Score is an editorial read of published capability, price and shipping pace. Nobody here has hands-on hours with Ollama. How we verify.
Worth it if
The strengths recorded against this entry.
- Local inference costs nothing per token and keeps prompts and files on the machine
- Works offline, which Ollama calls out specifically for mission-critical work
- Acts as the runtime under other agents, including ones named on its own front page, so it slots into existing stacks
- Cloud fallback covers models too large for a laptop without switching to a different tool or vendor
- States that prompt and response data is never logged or trained on, with zero data retention on the Team plan
Not worth it if
Any one of these blocks your use case.
- Cloud allowances are published only as multipliers ("50x more than Free", "5x more than Pro") against a base quantity Ollama never states
- The $100 Max tier is currently listed as paused for new sign-ups, so the top individual plan may not be available
- Local model quality is capped by your RAM and GPU, and open weights still trail frontier hosted models on the hardest tasks
- Comfort with a terminal, model files and quantisation choices is assumed, which puts it out of reach for non-technical users
What sets Ollama apart
- Local-first by default, so prompts and files stay on the machine unless a cloud model is chosen explicitly
- Positions itself as the runtime other agents plug into, naming Claude Code, OpenClaw and Codex on its own front page
- Cloud tier is an extension of the same tool rather than a separate product, with servers in the US, Europe and Singapore
- States that prompt and response data is never logged or trained on
- Free tier includes unlimited public models and light cloud usage alongside local inference
Key features
Local model execution
Runs open-weight models on your own hardware, which keeps prompts, code and documents on the machine and removes per-token cost entirely. Ollama lists offline capability for mission-critical work as an explicit reason to run this way.
Runtime for other tools
Ollama's own front page names Claude Code, OpenClaw and Codex as things you run against it. That positioning, as the layer under an agent rather than the agent itself, is why it shows up in so many stacks.
Optional cloud models
Larger models that will not fit on local hardware run on Ollama's cloud, served from the United States, Europe and Singapore. Free accounts get one concurrent cloud model, Pro three and Max ten.
Stated data policy
Ollama states that "Prompt or response data is never logged or trained on", and the Team plan is listed with zero data retention and logging. For anyone using the cloud path on work material, that policy is the load-bearing detail.
Cross-platform install
Downloads for macOS (macOS 14 Sonoma or later), Windows and Linux, with a shell-script install path on top of the direct downloads.
How it compares
| Tool | Best for | Pricing | Score |
|---|---|---|---|
| Ollama | Anyone who wants open-weight models running on their own hardware for privacy, offline work or zero marginal cost, with a cloud fallback for models that will not fit. | Free, with Pro at $20/mo and Team at $25/seat/mo for cloud usage | 8.8/10 |
| LM Studio vs LM Studio โ | People who want a graphical way to find, download and run open-weight models locally, plus a drop-in OpenAI-compatible endpoint for their own code. | Free desktop app; cloud inference is pay as you go per token | 8.5/10 |
| Osaurus | Developers on Apple Silicon Macs who want a private, local LLM server for coding assistants and automation without sending data to the cloud. | Free (open-source) | 7.5/10 |
| DeepSeek vs DeepSeek โ | Developers and teams running heavy coding or reasoning workloads who want frontier-level model performance without frontier-level API bills. | Free chat + API from $0.27/M input tokens | 8.9/10 |
| Aider vs Aider โ | Terminal-based developers who want an AI pair programmer that edits their local git repo directly, without adopting a new IDE. | Free (BYOK โ pay your LLM provider) | 8.6/10 |
Compare head-to-head
Related reading
Claude Eval Escapes: An 8-Day Disclosure Clock
Anthropic suspended the cyber evals July 23, notified affected organizations July 27, published July 31. What those dates settle and what they don't.
DeepSeek-V4-Flash API: Responses, Codex, Agent Setup
DeepSeek opened the V4-Flash public beta on July 31, 2026 with a native Responses API. Model ID, server-side state, Codex setup, and the agent claim.
Three Claude Eval Escapes in 141,000 Runs
Anthropic's July 30 disclosure came out of a retrospective sweep of 141,000+ cyber eval runs. The chronology, the arithmetic, and what it predicts.
Ready to try Ollama?
Head to the official site to start with Ollama โ pricing and plans are listed above.
Visit Ollama
