AI Solutions

Top Open-Weight LLMs You Can Self-Host in 2026

The top open-weight LLMs you can self-host in 2026: sizes, licences, hardware needs and serving tools like vLLM, SGLang, Ollama and llama.cpp.

GPTLabAI team 6 min read

The best open-weight LLMs to self-host in 2026 fall into two groups. Mid-sized models such as Qwen3.8-27B, Gemma 4, Devstral Small 2 and gpt-oss run on a single GPU or a well-specified workstation. Very large mixture-of-experts models such as DeepSeek-V4.1-Flash, GLM-5.3, Kimi K3 and Mistral Large 3 get close to frontier quality but need a multi-GPU server. Your licence requirements and hardware budget usually decide between them before benchmarks do.

Below we list the models worth shortlisting as of September 2026, explain how to estimate hardware, and compare the main serving tools.

Why self-host at all?

  • Data control — prompts and documents never leave your infrastructure, which simplifies GDPR, NDA and client security reviews.
  • Predictable cost — a fixed GPU bill instead of per-token pricing, attractive for steady, high-volume workloads.
  • Version stability — no forced upgrades or deprecations; you decide when to change models.
  • Customisation — fine-tuning, custom quantisation and full control over the serving stack.

The trade-offs are real: you run the GPUs, handle scaling and security patches, and the very best closed models still lead on the hardest tasks. Many teams run a hybrid setup. If you are weighing this up, our post on private RAG for business shows a typical self-hosted architecture.

Open-weight LLMs worth shortlisting (as of September 2026)

Parameter counts and licences below come from each model’s official Hugging Face card. Always read the full licence before commercial use.

Model Size / type Licence Realistic hardware Notes
gpt-oss-20b ~21B MoE Apache 2.0 ~16 GB memory Good reasoning for its size; laptops and small servers
gpt-oss-120b ~117B MoE (~5B active) Apache 2.0 Single 80 GB GPU Strong general model with low active compute
Gemma 4 E2B, E4B, 12B, 26B-A4B MoE, 31B dense Apache 2.0 Phones/laptops up to a single workstation GPU Multimodal, up to 256K context on larger sizes
Qwen3.8-27B 27B dense Apache 2.0 One 24–48 GB GPU when quantised 262K native context; strong all-rounder and coder
Devstral Small 2 24B Apache 2.0 Single RTX 4090 or 32 GB Mac (per model card) Built for agentic coding, 256K context
Llama 4 Scout 109B MoE (17B active) Llama 4 Community License Single H100 with int4 quantisation Older (2025) but widely supported; licence has conditions
Mistral Large 3 675B MoE (41B active) Apache 2.0 One 8-GPU node Permissive licence at the large end
DeepSeek-V4.1-Flash 552B MoE (small active set) MIT Multi-GPU server 1M context; served with vLLM or SGLang
GLM-5.3 753B Own GLM-5.3 licence Multi-GPU server Positioned by Z.ai as its strongest open coding model
Kimi K3 2.8T MoE (104B active) Own Kimi K3 licence Large multi-GPU cluster 1M context, multimodal; check licence conditions

A few selection notes from our projects:

  • Start in the 20–35B range. For internal assistants, document Q&A and coding help, a quantised model of this size on one GPU is often enough, especially with good retrieval.
  • Go large only when your evaluation says so. The jump from a 27B model to a 500B+ model is a jump from one GPU to a cluster.
  • “Open weights” is not the same as “open source”. Apache 2.0 and MIT are simple. Custom licences can add conditions for very large companies, attribution requirements or acceptable-use rules. Have someone read them.

How much hardware do you need?

A simple rule of thumb for the model weights:

Memory for weights ≈ parameters × bytes per parameter 16-bit ≈ 2 bytes, 8-bit ≈ 1 byte, 4-bit ≈ 0.5 bytes.

So a 27B model needs roughly 54 GB at 16-bit, 27 GB at 8-bit and 14 GB at 4-bit — before anything else. On top of that you need memory for the KV cache, which grows with context length and the number of simultaneous users. Long contexts and many concurrent requests can need as much memory as the weights themselves.

For mixture-of-experts (MoE) models, remember:

  • All parameters must fit in memory (GPU, or GPU plus CPU offload),
  • but only the active parameters are computed per token, so MoE models are faster than their total size suggests.

Practical tiers:

Setup What fits Typical use
Laptop / Mac with 16–32 GB ~4–24B models at 4-bit Local dev, prototyping, personal assistants
Single 24 GB consumer GPU Up to ~30B at 4-bit, short-to-medium context Small teams, internal tools
Single 80 GB data-centre GPU ~70–120B (quantised or MoE), or 30B with long context and many users Production for a department
8-GPU node Large MoE models (hundreds of billions of parameters) Company-wide, near-frontier quality

Quantisation to 8-bit usually costs little quality; 4-bit is usually acceptable for chat and retrieval but test it on your own tasks, especially for coding and maths.

Serving tools: vLLM, SGLang, Ollama, llama.cpp

Tool Best for Strengths Watch out for
vLLM Production GPU serving High throughput, continuous batching, OpenAI-compatible API, broad model support Needs GPUs and some tuning
SGLang Production GPU serving, large MoE models Fast serving, often first-day support for new large models Similar operational effort to vLLM
Ollama Local dev, small internal deployments One-command install, model library, simple API Not designed for high-concurrency production
llama.cpp CPUs, Macs, edge devices, mixed CPU/GPU Runs almost anywhere, GGUF quantisation, very efficient Lower throughput for many simultaneous users

Our usual path: prototype with Ollama or llama.cpp on a developer machine, then deploy with vLLM or SGLang behind an authenticated gateway for production. Because all of them can expose an OpenAI-compatible endpoint, application code rarely needs to change between stages.

Production checklist for self-hosted LLMs

  • Shortlist two or three models; evaluate on 50+ real tasks (see our LLM evaluation guide)
  • Read and record the licence terms for each model
  • Size GPU memory for weights plus KV cache at your target context and concurrency
  • Test the quantised version you will actually deploy, not the full-precision one
  • Put the model behind an API gateway with authentication, rate limits and logging
  • Keep the model endpoint off the public internet
  • Pin model and serving-engine versions; upgrade deliberately after re-testing
  • Monitor latency (p50/p95), throughput, GPU memory and error rates
  • Plan capacity: what happens when usage doubles?

Key takeaways

  • Mid-sized models (Qwen3.8-27B, Gemma 4, Devstral Small 2, gpt-oss) cover most business use cases on one GPU.
  • Large MoE models (DeepSeek-V4.1-Flash, GLM-5.3, Kimi K3, Mistral Large 3) approach frontier quality but need multi-GPU servers.
  • Licences range from Apache 2.0 and MIT to custom terms. Check before you ship.
  • Estimate memory as parameters × bytes per parameter, then add room for the KV cache.
  • Use Ollama or llama.cpp for local work, and vLLM or SGLang for production serving.

Getting a self-hosted model into production

Choosing the model is one decision among many: hardware sizing, quantisation, serving, security and integration with your apps all matter. Our LLM integration service covers the full path from shortlisting and evaluation to a secured, monitored deployment on your own infrastructure. If you are considering self-hosting, talk to us and we’ll help you size it realistically.

7 min

How Much Does an AI Chatbot for Business Cost?

What drives AI chatbot cost for a business: build, model API usage, hosting and maintenance, build vs buy, plus a simple formula to estimate your token costs.

Read article

6 min

Prompt Engineering Best Practices for Developers

Prompt engineering best practices for developers: clear instructions, examples, XML structure, reliable output formats, tool descriptions and eval-driven iteration.

Read article

Have a project in mind? Let’s talk.

Whether you run a business or a research group, tell us what you need built, fixed or evaluated. You get a free consultation and a clear written estimate — no obligation.

  • Free consultation
  • Written scope and estimate
  • We reply within one working day
Contact us