Skip to main content

Updated as of May, 2026. Calculated for Quantization Q4_K_M VRAM. Estimated throughput on an AMD EPYC CPU from Hetzner's CCX node cluster. Verify official requirements with Hugging Face before deployment.

Compatible models

NoneTemplateParameterMemory in the fourth quarter (Q4)BandwidthLicenseTuning is in progress...RiposoDanger or hazard that may result in an accident and/or damage to the system, data loss, financial losses for Ollama.Points
1Llama 3.2 3BLLaMA 3GGUF3B2 GB35 t/s LLaMA 3 qlora
82.0%
medium
65.6%
2Llama 3.1 8BLLaMA 3GGUF8B4.7 GB18 t/s LLaMA 3 qlora
92.0%
medium
65.2%
3Mistral 7B v0.3MistralGGUF7B4.1 GB20 t/s Apache 2.0 qlora
90.0%
low
65.0%
4Gemma 2 9BGemmaGGUF9B5.5 GB15 t/s Gemma lora
86.0%
medium
62.8%
5Aya 23 8BAyaGGUF8B4.7 GB17 t/s Apache 2.0 lora
77.0%
medium
60.5%

How is the score calculated?

  • Experience (30%)MTEB stability benchmarks and community support
  • Commercial License (20% )Commercial use allowed without restrictions
  • Bandwidth Utilization (30%)Tokens on the selected VPS (normalized to a maximum of 150 tokens/sec).
  • Match found in the codePrimary key (100% match) or secondary index (70% match) by Use Case

Come utilizzare SLM Picker

Choose your VPS size at Hetzner.

Choose CCX13/22/32/52: The tool automatically filters only compatible models with available RAM on that floor.

Set use case and target languages

Choose your primary use case (chat, RAG, code, translation...) and the European languages you need. No language selected = all templates.

Adjust with license and maximum size

Activate "Commercial Licenses Only" if you need commercial use without restrictions and limits maximum size to billions of parameters.

Sort and compare table rows by model

Click on Columns Param, Throughput, Maturity or Score to reorder the table and select the best model for your use case.

Suggerimenti

  • For commercial use, always activate the filter "Commercial Licenses Only" before selecting a model, even if the score seems high.
  • If your use case is RAG, prioritize models with primary match use case rather than maximum throughput.
  • Small VPS (CCX13) - prefer models with size ≤7B to maintain usable throughput in production.

Domande frequenti

How is the overall score for each model calculated?

It's a weighted average of four factors: maturity (30%, benchmark stability and community support), commercial license (20%), normalized throughput on the selected VPS (30%), and match with chosen use case (20%, 100% primary, 70% secondary).

What does the risk score (low/medium/high) indicate?

It's a composite risk indicator calculated from the model's catalog data (security, license, maturity): a "high" risk indicates an immature or licensable model to be evaluated with caution before deployment.

Does the throughput shown change based on the VPS I choose?

Yes, the column Throughput shows tokens at second estimated specifically for the selected Hetzner CCX size (CCX13/22/32/52): changing VPS the table is recalculated.

What happens if I don't select any target language?

Language filter is optional: if you select no language, the tool will show all available models regardless of supported languages.

Are VRAM and throughput data reliable for a production deployment?

Estimated updated as of May 2026, calculated for Q4_K_M quantization on AMD EPYC CPU (Hetzner CCX): before deployment, always verify official model requirements on HuggingFace.