Ollama changes pricing for the Pro, Max and Team plans to transparent per-token pricing
Ollama has introduced transparent per-token pricing for the Pro, Max and Team plans, with monthly credit replacing fixed 5-hour and weekly limits. A shared Team plan is also newly available at $500/month for an unlimited number of users.
From 31 August 2026, Ollama is changing pricing for the Pro, Max and Team plans from the previous GPU-time-based model to transparent per-token pricing. Each plan now includes a monthly credit pool, drawn down at the per-token rate published on the pricing page and on the page for each model; once the credit is exhausted, usage can continue at the same rate. The Pro plan costs $20 per month and includes $60 of monthly usage, while Max costs $100 per month with $300 of usage. A Team plan is also newly available at an introductory price of $500 per month, offering a shared pool of $1 000 in monthly usage for an unlimited number of users with no limits on individual accounts, and an overview of usage across the entire organization in one place.
According to the company, the new plans have no service fees or hidden limits and remove the previous 5-hour and weekly limits of the old model. The monthly credit resets each month and does not roll over to the next period. The service runs on dedicated compute resources in the USA and Europe, and also in Singapore for a limited set of Qwen models, with stated zero data retention — the company says it does not log prompts or train on user data. The plans work with Claude Code and Codex and also offer their own API. The free plan now includes a small amount of monthly usage for selected starter models, with the option to top up credit and switch to pay-as-you-go without a subscription.
Existing subscribers to the Pro, Max and Team plans remain on their current pricing and can switch to the new pricing at any time in their account settings; switching resets usage and ends the old limit model, while the monthly credit renewal date remains unchanged. The company explains the change by saying that GPU-time-based pricing was difficult to predict, especially as open models grew in size — it cites Kimi K3 with 2.8 trillion parameters as an example.
Why it matters
For individual users of the Pro and Max plans, the change means an end to fixed time limits, but also a need to track actual token usage within the monthly credit. For companies, the new Team plan offers a more predictable shared budget without per-seat fees for a team of any size, making it easier to plan the cost of running open models compared with the previous GPU-time-based pricing.
Two audiences, two different impacts
What this means
For individuals
Individual users of the Pro and Max plans are no longer subject to 5-hour or weekly limits, but usage is now calculated by tokens within the monthly credit, so actual spending on individual requests needs to be monitored.
For a business
Companies deploying open models through Ollama gain a new Team plan with shared credit for an unlimited number of users without per-seat fees, changing team cost calculations compared with earlier plans with fixed limits.
DevelopmentCheck the original
Event sources
clearly official source · 1 publisher, 0 independent. We count feeds from the same owner only once.