Skip to main content

Disclaimer: prices are manual snapshots (2026-05) and may vary. Hetzner costs are indicative. This tool is not financial advice. Always check the official price lists of providers before making a purchase decision.

Configure the workload

Average number of API calls to the LLM model per day
Include the system prompt and context. Use the Token Counter to estimate.
0% = constant load, 10% doubles around every 7 months
Approximate exchange rate: USD/EUR approximately $1 = €0.92 (client side)
Selected server is Hetzner.CCX13(€21.16/mese fisso)

Cost Comparison Tool

Model/ ProviderMonthly cost (month 1)Time to Complete: 12 monthsOllama vs Federico Calò's solution (on premises)Breakeven point
Llama 3.1 8B (Ollama on-prem)Meta / Self-hostedOn-Premises€21.16€253.92Reference-
DeepSeek V3DeepSeekCheaper€12.83€274.45 +€20.53 Month 11
Gemini 2.5 ProGoogle€58.65€1254.19 +€1000.27 Month 1
Mistral Large 2Mistral AI€77.28€1652.58 +€1398.66 Month 1
GPT-4oOpenAI€117.30€2508.38 +€2254.46 Month 1
Claude Sonnet 4.6Anthropic€165.60€3541.24 +€3287.32 Month 1
GPT-5OpenAI€234.60€5016.75 +€4762.83 Month 1
GPT-5 ProOpenAI€703.80€15050.26 +€14796.34 Month 1
Claude Opus 4.7Anthropic€869.40€18591.50 +€18337.58 Month 1

Cumulative TCO - Cloud vs Ollama On-Prem

Most cost-effective reference as a Cloud model.

Decision

Cheaper cloud

Break even reached in the month of , on the most economical cloud plan, at 1%__. with a total cost reduction of 2%.1

Infrastructure cost Ollama (Hetzner CCX13): fixed €21.16/month. Independent of volume.

Come utilizzare LLM Cost Optimizer

Describe your workload

Insert query/day, input/output and growth rate tokens. Load a predefined high-volume scenario.

Choose currency and horizon

Select EUR or USD and simulation horizon (1-36 months) to project cumulative TCO over time.

Compare Cloud vs On-Prem Ollama

Read the table with 8 cloud models + Ollama scenarios on-prem Hetzner, automatic verdict and break-even month.

Suggerimenti

  • Use the "high volume" scenario to quickly stress test when it's convenient to switch to Ollama on-prem.
  • If the break-even point is beyond the simulated horizon, extend the horizon up to 36 months to see if it changes the verdict.
  • Combine this tool with the Token Counter Multi-Model to estimate your real workload's input/output token averages with high accuracy.

Domande frequenti

How is break-even point calculated with Ollama on-prem?

Cloud cost vs. infrastructure fixed cost, with a break-even point where cumulative cloud costs exceed on-premises costs for the first time.

What is included in the cost of the Ollama on-prem scenario?

Monthly fixed cost of Hetzner VPS (CCX13 or CCX23 tier) only, independent of query volume. No setup time, maintenance costs, or additional bandwidth fees included.

Why is the verdict sometimes "borderline"?

The verdict is borderline when the break-even point falls near the end of the simulated horizon: in these cases, the result depends heavily on assumptions about growth and volume, so it should be interpreted with caution.

Are prices for cloud models automatically updated?

No, prices for our 8 cloud models are reference values inserted into the tool: always check the official pricing list of the provider before making budget decisions.

Does calculation take into account growth of query volume over time?

Yes, the monthly percentage growth set will be compounded month over month, both for Cloud costs and compared to our on-premises infrastructure at fixed costs.