Back
Kevin Riedl

9 min read · 25 Jun 2026
Last reviewed

Next
Made on your device, with no Instagram connection. We copy the post link for Instagram’s Link sticker.

Open-Weight LLMs 2026: DeepSeek, Qwen, Kimi, GLM, Llama and Muse

There is no universal best open-weight model. This shortlist was reviewed on 5 October 2026. It covers DeepSeek V4, the very different Qwen3.8 size and license tiers, Kimi K3, GLM-5.3, Llama 4 and Meta’s Muse Glimmer-30B. For workstation-scale agent evaluations, compare Qwen3.8-27B with Muse Glimmer-30B. Do not transfer Llama 4’s EU multimodal developer restriction to Muse’s separate Apache 2.0 release.

"Open weight" only says that weights are downloadable. It does not guarantee an Open Source Initiative-approved license, low hardware cost, identical hosted and self-hosted behavior, or usable quality at the advertised maximum context. Shortlist by the exact checkpoint and license, then test your own tasks.

This is a reviewed buyer shortlist, not an exhaustive claim to list every latest model. We checked the cited checkpoints and license sources on 5 October 2026 and added the previously omitted Muse Glimmer release. Vendor benchmark tables document their own evaluation setups, not a neutral cross-vendor ranking or Wavect-run test.

Need a model shortlist that fits your tasks, EU boundary, and GPU budget?

 Shortlist Models for My Workload

What separates the current open-weight families?

  • DeepSeek V4. The current official repositories include V4 Flash-0731 and V4 Pro-0813. Both are MIT-licensed. The original V4 release card documents 1M-token context and sparse architectures, while the dated checkpoints supersede the preview releases.
  • Qwen3.8. Alibaba's official Qwen3.8 collection includes a 27B multimodal model and a 2.4T total, 95B-active sparse model. The 27B card states a native 262,144-token window, extension to 1M with YaRN, and Apache 2.0 licensing. The 2.4T checkpoint uses the Qwen3.8-Max License, including conditions for very large products and large MaaS or AI work-assistant businesses. Extension is a serving configuration, not evidence that every task remains reliable at 1M.
  • Kimi K3. Moonshot's model card states 2.8T total and 104B active parameters, native text, image, and video input, and a 1M-token context. The Kimi K3 License adds conditions for large model-as-a-service businesses and very large commercial products, while exempting internal use from those extra conditions.
  • GLM-5.3. Z.ai released GLM-5.3 after GLM-5.2. It uses the same 753B base model, supports 1M-token evaluation, and changes the post-training for coding and long-horizon work. Its GLM-5.3 License is permissive, but it is not labelled MIT, so the earlier MIT claim is no longer accurate for the current checkpoint.
  • Llama 4. Meta's model card lists Scout at 109B total and 17B active parameters with 10M context, and Maverick at 400B total and 17B active with 1M context. The use policy withholds the multimodal license grant from EU-domiciled individuals and companies whose principal place of business is in the EU. It expressly says that this restriction does not apply to end users of a product incorporating the models.
  • Muse Glimmer-30B. Meta’s official model card identifies an August 2026 release under Apache 2.0, approximately 29.6B total parameters including its perception encoder, text and image input, text output and a published 131,072+ context. It targets local agent workflows. This is a separate model and license from Llama 4, not a newly removed restriction on Llama itself.

Germany's Soofi S sits in a different category: a sovereign German-English project. Our Soofi S buyer's review separates public evidence from release plans. Our Kimi K3 API review for EU companies treats Moonshot's hosted service as a separate procurement decision from running the downloadable weights.

Size, context, license, and deployment

FamilyRepresentative current checkpointPublished sizePublished contextWeight licenseDeployment implication
DeepSeekV4 Flash-0731 / Pro-0813304B / 1.7T totalV4 family: 1MMITLarge multi-GPU deployment; calculate from the exact weight format
QwenQwen3.8-27B / 2.4T-A95B27B dense / 2.4T total, 95B active27B: 262,144 native, extensible to 1MApache 2.0 / Qwen3.8-Max License27B is materially easier to host; the 2.4T tier is cluster-scale
KimiKimi K32.8T total / 104B active1MKimi K3 LicenseCluster-scale checkpoint plus model-specific license review
GLMGLM-5.3753B total1MGLM-5.3 LicenseSerious multi-GPU or offloaded experimental serving
LlamaLlama 4 Scout / Maverick109B / 17B active; 400B / 17B active10M / 1MLlama 4 Community LicenseResolve the EU multimodal developer restriction first
MuseMuse Glimmer-30BApproximately 29.6B total, dense131,072+Apache 2.0Local-agent candidate; size the quantization, encoder and KV cache together

These are maximum context claims, not quality guarantees. Long sequences increase KV-cache demand and can reduce concurrency. Quantization changes memory needs and sometimes output quality. Do not convert parameter counts into a fixed GPU count without specifying precision, runtime, context, batch size, and redundancy.

Separate workstation candidates from cluster deployments

Qwen3.8-27B and Muse Glimmer-30B belong in a different deployment conversation from trillion-parameter checkpoints. Meta gives 24 GB and 32 GB target configurations for its quantized Muse packages, and 64 GB for full precision. Those are vendor targets for specified configurations, not a promise that every runtime, context length or concurrent workload will fit. Preserve headroom for KV cache, images, any drafter and the rest of the application. Our Muse Glimmer local-agent guide covers the specific packaging and tests rather than duplicating that review here.

Which family should you test for coding and agents?

The primary sources support a shortlist, not a universal ranking. DeepSeek's Flash-0731 card, Qwen3.8's model card, Kimi K3's card, and GLM-5.3's card all publish coding and agent benchmarks, but they use different agent harnesses, reasoning settings, task versions, and sometimes internal test sets. Compare numbers only where the task, split, harness, tool access, and scoring method match.

  • For a workstation-scale multimodal agent: evaluate Qwen3.8-27B and Muse Glimmer-30B side by side. Both use Apache 2.0 for these exact checkpoints, but context, tool-use behavior, quantization and runtime support differ.
  • For large coding or tool-use evaluations: test DeepSeek V4 Flash-0731, GLM-5.3, and Kimi K3 on completed-task rate, recovery behavior, latency, and total serving cost.
  • For very long context: test retrieval accuracy and task completion at your actual lengths. A 1M or 10M limit only establishes accepted input size.
  • For Llama 4 in the EU: obtain legal review before a developer downloads or deploys the multimodal weights. Do not misread the end-user exception as a developer license.
Kevin Riedl

"A model card tells you what to test and under which license. Your own eval tells you what to ship. Keep those two decisions separate."

Can you self-host these in the EU?

DeepSeek V4, Qwen3.8-27B and Muse Glimmer-30B use standard permissive licenses for the named weights. Qwen3.8-2.4T-A95B, GLM-5.3 and Kimi K3 use model-specific texts that require separate review. Llama 4’s multimodal policy withholds developer rights for the specified EU-based users while excepting end users of an incorporating product. That Llama-specific restriction is not Muse’s license. No weight license by itself establishes GDPR compliance.

Self-hosting can keep inference inside an EU-controlled environment, but it does not automatically establish GDPR compliance. You still need a lawful basis, data minimization, retention rules, access controls, processor and transfer analysis where relevant, logging policy, and an incident process. A hosted endpoint must be assessed separately from the downloadable checkpoint because its data location and contract are service-specific.

A defensible selection process

  1. Name the exact checkpoint. Family labels hide material differences in weights, modalities, context configuration, and license.
  2. Clear license and jurisdiction. Record the license revision and have counsel review model-specific restrictions when the deployment is consequential.
  3. Budget the actual serving configuration. Include precision, context, concurrency, runtime overhead, redundancy, and network topology.
  4. Run the same private eval. Use representative tasks, stable prompts, identical tool access, and explicit pass conditions across at least two candidates.
  5. Pin and monitor. Record the weight revision, tokenizer, runtime, quantization, and safety configuration. Re-evaluate before changing any of them.

Why the eval harness matters

Public benchmarks can suffer from contamination, overfitting, and harness effects. More importantly, they are not your workload. A few dozen representative tasks with clear pass criteria can expose failure modes that an aggregate leaderboard hides. Measure completed-task quality, latency, cost, tool errors, long-context retrieval, and recovery after a failed step.

Frequently Asked Questions

Does Meta’s Llama 4 EU restriction apply to Muse Glimmer?
Do not conflate the releases. The reviewed Muse Glimmer-30B card specifies Apache 2.0. Llama 4’s multimodal use policy contains its own EU developer restriction and end-user exception. Review the exact checkpoint and applicable law, rather than assigning one license to all Meta models.
Which models should a small team test locally first?
In this shortlist, Qwen3.8-27B and Muse Glimmer-30B are the workstation-scale candidates. Confirm the actual weight package, runtime, available memory, context length and concurrency before treating either as a fit.
Is the longest advertised context the best model?
No. Maximum accepted input size does not establish retrieval accuracy, task completion quality or affordable concurrent serving. Test realistic lengths and failure recovery.
Are these Wavect benchmark results?
No. Model specifications and release benchmarks come from the linked primary sources. The selection process is our engineering recommendation; it is not a claim that Wavect ran a new cross-model benchmark.

Final thoughts

The October 2026 shortlist now distinguishes workstation-scale Qwen3.8-27B and Muse Glimmer-30B from the much larger checkpoints. Keep license decisions checkpoint-specific, particularly for Muse versus Llama 4. Choose the smallest legally deployable configuration that passes your own acceptance tests, and assess hosted APIs separately from downloadable weights.

Production AI help

Building an AI product and worried about inference cost, architecture, or production readiness? Wavect helps founders turn AI prototypes into reliable production systems.

Explore the service path:

Inbox, without the noise

Follow the work that matters to you

Get a short email when we publish something new. Follow the whole blog or only the problems you care about.

What would you like to receive?
Choose your topics

Free, double opt-in, no tracking pixels.

Back
Kevin Riedl

9 min read · 25 Jun 2026
Last reviewed

Next

Get the next AI and agents field note

One concise email when we publish. No tracking pixels, and no inbox filler.

Free, double opt-in, no tracking pixels.