Top Open-Weight LLMs You Can Self-Host in 2026
The top open-weight LLMs you can self-host in 2026: sizes, licences, hardware needs and serving tools like vLLM, SGLang, Ollama and llama.cpp.

The best open-weight LLMs to self-host in 2026 fall into two groups. Mid-sized models such as Qwen3.8-27B, Gemma 4, Devstral Small 2 and gpt-oss run on a single GPU or a well-specified workstation. Very large mixture-of-experts models such as DeepSeek-V4.1-Flash, GLM-5.3, Kimi K3 and Mistral Large 3 get close to frontier quality but need a multi-GPU server. Your licence requirements and hardware budget usually decide between them before benchmarks do.
Below we list the models worth shortlisting as of September 2026, explain how to estimate hardware, and compare the main serving tools.
Why self-host at all?
- Data control — prompts and documents never leave your infrastructure, which simplifies GDPR, NDA and client security reviews.
- Predictable cost — a fixed GPU bill instead of per-token pricing, attractive for steady, high-volume workloads.
- Version stability — no forced upgrades or deprecations; you decide when to change models.
- Customisation — fine-tuning, custom quantisation and full control over the serving stack.
The trade-offs are real: you run the GPUs, handle scaling and security patches, and the very best closed models still lead on the hardest tasks. Many teams run a hybrid setup. If you are weighing this up, our post on private RAG for business shows a typical self-hosted architecture.
Open-weight LLMs worth shortlisting (as of September 2026)
Parameter counts and licences below come from each model’s official Hugging Face card. Always read the full licence before commercial use.
| Model | Size / type | Licence | Realistic hardware | Notes |
|---|---|---|---|---|
| gpt-oss-20b | ~21B MoE | Apache 2.0 | ~16 GB memory | Good reasoning for its size; laptops and small servers |
| gpt-oss-120b | ~117B MoE (~5B active) | Apache 2.0 | Single 80 GB GPU | Strong general model with low active compute |
| Gemma 4 | E2B, E4B, 12B, 26B-A4B MoE, 31B dense | Apache 2.0 | Phones/laptops up to a single workstation GPU | Multimodal, up to 256K context on larger sizes |
| Qwen3.8-27B | 27B dense | Apache 2.0 | One 24–48 GB GPU when quantised | 262K native context; strong all-rounder and coder |
| Devstral Small 2 | 24B | Apache 2.0 | Single RTX 4090 or 32 GB Mac (per model card) | Built for agentic coding, 256K context |
| Llama 4 Scout | 109B MoE (17B active) | Llama 4 Community License | Single H100 with int4 quantisation | Older (2025) but widely supported; licence has conditions |
| Mistral Large 3 | 675B MoE (41B active) | Apache 2.0 | One 8-GPU node | Permissive licence at the large end |
| DeepSeek-V4.1-Flash | 552B MoE (small active set) | MIT | Multi-GPU server | 1M context; served with vLLM or SGLang |
| GLM-5.3 | 753B | Own GLM-5.3 licence | Multi-GPU server | Positioned by Z.ai as its strongest open coding model |
| Kimi K3 | 2.8T MoE (104B active) | Own Kimi K3 licence | Large multi-GPU cluster | 1M context, multimodal; check licence conditions |
A few selection notes from our projects:
- Start in the 20–35B range. For internal assistants, document Q&A and coding help, a quantised model of this size on one GPU is often enough, especially with good retrieval.
- Go large only when your evaluation says so. The jump from a 27B model to a 500B+ model is a jump from one GPU to a cluster.
- “Open weights” is not the same as “open source”. Apache 2.0 and MIT are simple. Custom licences can add conditions for very large companies, attribution requirements or acceptable-use rules. Have someone read them.
How much hardware do you need?
A simple rule of thumb for the model weights:
Memory for weights ≈ parameters × bytes per parameter 16-bit ≈ 2 bytes, 8-bit ≈ 1 byte, 4-bit ≈ 0.5 bytes.
So a 27B model needs roughly 54 GB at 16-bit, 27 GB at 8-bit and 14 GB at 4-bit — before anything else. On top of that you need memory for the KV cache, which grows with context length and the number of simultaneous users. Long contexts and many concurrent requests can need as much memory as the weights themselves.
For mixture-of-experts (MoE) models, remember:
- All parameters must fit in memory (GPU, or GPU plus CPU offload),
- but only the active parameters are computed per token, so MoE models are faster than their total size suggests.
Practical tiers:
| Setup | What fits | Typical use |
|---|---|---|
| Laptop / Mac with 16–32 GB | ~4–24B models at 4-bit | Local dev, prototyping, personal assistants |
| Single 24 GB consumer GPU | Up to ~30B at 4-bit, short-to-medium context | Small teams, internal tools |
| Single 80 GB data-centre GPU | ~70–120B (quantised or MoE), or 30B with long context and many users | Production for a department |
| 8-GPU node | Large MoE models (hundreds of billions of parameters) | Company-wide, near-frontier quality |
Quantisation to 8-bit usually costs little quality; 4-bit is usually acceptable for chat and retrieval but test it on your own tasks, especially for coding and maths.
Serving tools: vLLM, SGLang, Ollama, llama.cpp
| Tool | Best for | Strengths | Watch out for |
|---|---|---|---|
| vLLM | Production GPU serving | High throughput, continuous batching, OpenAI-compatible API, broad model support | Needs GPUs and some tuning |
| SGLang | Production GPU serving, large MoE models | Fast serving, often first-day support for new large models | Similar operational effort to vLLM |
| Ollama | Local dev, small internal deployments | One-command install, model library, simple API | Not designed for high-concurrency production |
| llama.cpp | CPUs, Macs, edge devices, mixed CPU/GPU | Runs almost anywhere, GGUF quantisation, very efficient | Lower throughput for many simultaneous users |
Our usual path: prototype with Ollama or llama.cpp on a developer machine, then deploy with vLLM or SGLang behind an authenticated gateway for production. Because all of them can expose an OpenAI-compatible endpoint, application code rarely needs to change between stages.
Production checklist for self-hosted LLMs
- Shortlist two or three models; evaluate on 50+ real tasks (see our LLM evaluation guide)
- Read and record the licence terms for each model
- Size GPU memory for weights plus KV cache at your target context and concurrency
- Test the quantised version you will actually deploy, not the full-precision one
- Put the model behind an API gateway with authentication, rate limits and logging
- Keep the model endpoint off the public internet
- Pin model and serving-engine versions; upgrade deliberately after re-testing
- Monitor latency (p50/p95), throughput, GPU memory and error rates
- Plan capacity: what happens when usage doubles?
Key takeaways
- Mid-sized models (Qwen3.8-27B, Gemma 4, Devstral Small 2, gpt-oss) cover most business use cases on one GPU.
- Large MoE models (DeepSeek-V4.1-Flash, GLM-5.3, Kimi K3, Mistral Large 3) approach frontier quality but need multi-GPU servers.
- Licences range from Apache 2.0 and MIT to custom terms. Check before you ship.
- Estimate memory as parameters × bytes per parameter, then add room for the KV cache.
- Use Ollama or llama.cpp for local work, and vLLM or SGLang for production serving.
Getting a self-hosted model into production
Choosing the model is one decision among many: hardware sizing, quantisation, serving, security and integration with your apps all matter. Our LLM integration service covers the full path from shortlisting and evaluation to a secured, monitored deployment on your own infrastructure. If you are considering self-hosting, talk to us and we’ll help you size it realistically.