Open AI Models 2026: Qwen 3.8, DeepSeek V4, GLM 5.3 and Nemotron
As of August 2026, practical local candidates include Qwen3.8-27B, Nemotron Lightning, Muse Glimmer, Gemma 4 and TAIDE. This guide compares licensing, hardware, agent cost and data sovereignty.

Updated August 15, 2026: Enterprise model selection has moved far beyond the Llama 3, Qwen 2.5 and DeepSeek R1 landscape. The current shortlist includes Qwen 3.8, DeepSeek V4-Flash-0731, GLM 5.3, Kimi K3, NVIDIA Nemotron 3.5 Lightning, Meta Muse Glimmer 30B, Gemma 4 and Taiwan's TAIDE. This guide separates hosted services from downloadable weights and compares licenses, infrastructure, agent reliability, multilingual performance and data residency.
Executive answer: Qwen3.8-27B, Nemotron 3.5 Lightning, Muse Glimmer 30B, Gemma 4 and Gemma-3-TAIDE-12B are the most practical first-round on-premises candidates. DeepSeek V4-Flash-0731 is the cost-focused agent reference; Qwen3.8-2.4T, Kimi K3 and Nemotron 3 Ultra require data-center-class infrastructure. GLM 5.3 is available through Coding Plan, with its model API announced as forthcoming, but should not yet be described as a newly released downloadable checkpoint.
“Open source” can mean three different things
- Hosted API: fast to adopt, but data leaves your environment and the provider controls pricing, versions and retirement schedules.
- Open-weight model: weights can be downloaded and usually fine-tuned, while training data and the complete production process may remain unavailable.
- Open-source AI: the Open Source AI Definition 1.0 requires more than downloadable parameters: users need the information and code needed to study, modify and recreate the system.
This distinction is operational, not semantic. A hosted model is not a candidate when inference data must remain inside a private network. A downloadable model may still be unsuitable when the license restricts redistribution, derivative services or specific commercial uses.
2026 model landscape: specifications that affect deployment
| Model | Access | Published specification | Context | Enterprise takeaway |
|---|---|---|---|---|
| Qwen3.8-27B | Open weights; API forthcoming | 27B dense, native vision-language | 262K native; extendable to 1M | Apache 2.0; a practical local Qwen baseline with image, video and adjustable reasoning. |
| Qwen3.8-2.4T-A95B / Max | API and open weights | 2.4T total / 95B active | 262K native; about 1.01M extended | Uses the custom qwen3.8-max license and needs a large cluster. Do not apply the 27B license to this checkpoint. |
| DeepSeek V4-Flash-0731 | API and open weights | 284B / 13B active, with DSpark | 1M; up to 384K output | MIT; optimized for agent task economics, coding and tool use. |
| DeepSeek V4-Pro-0813 | API and open weights | About 1.7T total / 49B active | 1M; up to 384K output | MIT. Higher-complexity reasoning and long-running agents; it became the version behind the deepseek-v4-pro endpoint on August 12. Weights are downloadable, but full self-hosting is far heavier than Flash. |
| GLM 5.3 | Coding Plan; API forthcoming | GLM 5.2 base with post-training upgrades | 1M; up to 128K output | Text-only, always-thinking model with low/high/max effort. It is currently a service release, not a new open-weight release. |
| Kimi K3 | API and open weights | 2.8T / 104B active | 1,048,576 tokens | Native multimodality and long-horizon knowledge work, with a custom Kimi K3 license and data-center deployment needs. |
| Nemotron 3.5 Lightning | API/NIM and open weights | 30B / 3B active | 1M | OpenMDW-1.1, NVFP4 and BF16; designed for high-throughput, low-latency agent execution. |
| Nemotron 3 Ultra | API/NIM plus weights, data and recipes | 550B / 55B active | 1M | A data-center reasoning and coordination model in NVIDIA's broader open deployment stack. |
| Meta Muse Glimmer 30B | Open weights | About 29.6B dense plus vision encoder | 131K+ | Apache 2.0; official 4-bit releases target local multimodal agents on 24GB/32GB devices. |
| Gemma 4 | API and open weights | E2B, E4B, 26B MoE (about 3.8B active) and 31B Dense | 128K edge / up to 256K larger models | Apache 2.0, image and video support throughout, plus audio on E2B and E4B. |
Specifications and licenses are based on official release pages as of the update date. Maximum context does not guarantee equal accuracy, speed or cost at that length.
Enterprise selection matrix
This is an editorial 1–5 screening aid derived from official specifications, supported modalities, model size and the enterprise use cases in this guide. It is not a synthetic cross-vendor benchmark.
| Model | Agent / coding | Multimodal | Traditional Chinese potential | Local deployability | Suggested role |
|---|---|---|---|---|---|
| Qwen3.8-27B | ★★★★★ | ★★★★★ | ★★★★★ | ★★★★★ | Chinese multimodal baseline |
| DeepSeek V4-Flash-0731 | ★★★★★ | ★★★★★ | ★★★★★ | ★★★★★ | Cost-efficient agent ceiling |
| GLM 5.3 | ★★★★★ | ★★★★★ | ★★★★★ | Service only | Engineering-agent reference |
| Kimi K3 | ★★★★★ | ★★★★★ | ★★★★★ | ★★★★★ | Long-context multimodal ceiling |
| Nemotron 3.5 Lightning | ★★★★★ | ★★★★★ | ★★★★★ | ★★★★★ | High-throughput execution layer |
| Muse Glimmer 30B | ★★★★★ | ★★★★★ | ★★★★★ | ★★★★★ | Local multimodal agent |
| Gemma 4 E4B | ★★★★★ | ★★★★★ | ★★★★★ | ★★★★★ | Edge multimodality |
| Gemma-3-TAIDE-12B | ★★★★★ | ★★★★★ | ★★★★★ | ★★★★★ | Taiwan-language baseline |
How to read it: five stars means a strong fit for this guide's enterprise scenario, not universal model superiority. Re-test Traditional Chinese with your own Taiwan-specific evaluation set.
How the leading model families differ
Qwen 3.8: a deployable 27B model and a cluster-scale 2.4T flagship
Qwen3.8-27B is an Apache 2.0 dense vision-language model for text, images and video. It supports a 262,144-token native context, extension to 1M and adjustable reasoning. It is the practical on-premises Qwen candidate.
Qwen3.8-2.4T-A95B exposes Max-class weights but uses a custom license and requires data-center infrastructure.
DeepSeek V4-Flash-0731: optimize completed-task cost, not token price alone
V4-Flash-0731 is the primary DeepSeek version in this guide. It retrains the 284B/13B-active architecture for stronger agent, coding and tool performance and adds DSpark speculative decoding. Its weights use the MIT license.
Agent economics means measuring turns, retries and total tokens required to complete a task. On the update date the official API listed cache-miss input/output at US$0.14/0.28 per million tokens; note that DeepSeek has announced peak and off-peak pricing from August 16, 2026, with off-peak at half the peak rate, so factor in when your workload actually runs. Industry analysis frames this as task-completion economics.
GLM 5.3: a post-training service upgrade
GLM 5.3 is fully available in GLM Coding Plan, while its model API is marked forthcoming. It reuses the GLM 5.2 base and improves complex engineering, terminal agents and security tasks through post-training. It is text-only, has 1M context, up to 128K output and always-on reasoning with low/high/max effort. Unlike GLM 5.2, which shipped MIT weights on Hugging Face almost immediately, the vendor says GLM 5.3 weights and broader API access will be released in stages after safety evaluation, so keep it out of on-premises capacity planning until they actually land.
Kimi K3: native multimodality and a one-million-token context
Kimi K3 is a 2.8T/104B-active multimodal MoE for long-horizon coding, research and knowledge work. Its custom license and enormous infrastructure requirement need separate legal and engineering reviews.
NVIDIA Nemotron: Lightning for throughput, Ultra for difficult coordination
Nemotron 3.5 Lightning 30B-A3B provides NVFP4 and BF16 weights for high-frequency tool calls and low-latency agent workloads. Nemotron 3 Ultra is a 550B/55B-active coordinator with open weights, data and training recipes. NVIDIA's strategy is a system of Lightning, Super, Ultra, retrieval, safety and speech models deployed through NIM, TensorRT-LLM, Dynamo and NeMo.
Meta Muse Glimmer 30B: a local multimodal agent for consumer hardware
Muse Glimmer 30B combines a roughly 29.6B dense language model with a vision encoder, 131K+ context, tool use, planning and recovery. It uses Apache 2.0. Meta supplies full-precision, two 4-bit variants and a DFlash drafter; the quantized language model is under 20GB and targets complete operation within 24GB/32GB memory budgets.
Smaller models are often the better production models
- Qwen3.8-27B: balanced Chinese, vision, coding and agent capability.
- Nemotron 3.5 Lightning: strong throughput and NVIDIA deployment integration.
- Muse Glimmer 30B: local image-and-text agents on consumer devices.
- Gemma 4 E2B / E4B / 26B MoE / 31B: edge-to-workstation coverage and a Western supply-chain option.
- gpt-oss-20b / 120b: Apache 2.0 text reasoning with tool and structured-output strengths.
- Ministral 3: compact 3B, 8B and 14B options for edge and multilingual use.
- Gemma-3-TAIDE-12B-Chat-2602: Taiwan text and Traditional Chinese office workflows.
Gemma 4: one family from edge devices to workstations
Gemma 4 uses Apache 2.0 and comes in four sizes: E2B, E4B, 26B MoE and 31B Dense. Every model accepts images and video; E2B and E4B also accept audio. The edge models use 128K context, while the larger models offer up to 256K. Start with E4B for offline mobile and IoT workflows, then test quantized 26B or 31B for workstation reasoning. The 26B model activates about 3.8B parameters for latency, while 31B Dense prioritizes quality and fine-tuning flexibility.
Why TAIDE still belongs in a Taiwan evaluation
Gemma-3-TAIDE-12B-Chat-2602 is based on Gemma 3, not Gemma 4. Its value is Taiwan language, office tasks and local usage rather than global English rankings. Test agency names, Taiwan laws, Minguo/Gregorian dates, addresses, currency, full-width punctuation and cross-strait vocabulary in the same suite used for Qwen, Gemma, Muse and Nemotron.
Hardware planning: active parameters are not memory requirements
| Environment | Reasonable range | Typical use | Warning |
|---|---|---|---|
| 16–32GB | Quantized 3B–12B; selected MoE models | Individual document work and offline summaries | Long contexts quickly expand KV cache. |
| 32–64GB | Quantized 12B–35B | Department PoC and low-concurrency RAG | Qwen3.8-27B, Muse Glimmer and Nemotron Lightning still require workload tests. |
| 80GB single/multi-GPU | Quantized 70B–120B or MoE | Shared enterprise services | Measure prefill, decode, tensor parallelism and failover. |
| Multi-node cluster | Qwen 2.4T, DeepSeek V4, Kimi K3, Nemotron Ultra | Frontier quality and high-scale service | Interconnects and operations dominate the project. |

LM Studio, Ollama, vLLM and SGLang serve different stages
LM Studio and Ollama make local experiments easy. vLLM and SGLang address multi-user throughput, long context and production serving. None supplies enterprise identity, authorization, rate limits, audit logs, content safety, version pinning and failover by itself.
Build your own evaluation set
Use 100–300 representative tasks and measure accuracy, citation quality, extraction F1, test pass rate, Traditional Chinese, tool/schema reliability, P50/P95 latency, prompt injection resistance and one-year TCO. Repeat each item at least three times and preserve the model ID, weight hash, quantization, inference engine and sampling settings.
China and the United States are pursuing different open-model strategies
The simplistic “China is open, America is closed” story no longer works. Chinese developers compete through very large MoE models, one-million-token contexts, frequent date-stamped post-training releases and aggressive API pricing. Qwen, DeepSeek, GLM and Kimi focus on coding, tools and long-running agents, but their licenses range from Apache 2.0 and MIT to custom terms.
The US ecosystem remains dual-track: frontier proprietary APIs coexist with open-weight families built around deployable sizes, permissive licenses, edge/data-center hardware and complete tool, data and safety stacks. Gemma 4, Muse Glimmer, Nemotron and gpt-oss illustrate that route. The distinction is an ecosystem strategy, not a universal capability ranking.

Jensen Huang's first post on X: open models strengthen safety, innovation and sovereignty
On July 24, 2026, NVIDIA CEO Jensen Huang used his first post on X to share a letter NVIDIA had signed on why open models matter. His argument was direct: AI will transform every industry, power every company and be built by every country, while open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The post linked to the joint Open Weights and American AI Leadership letter. NVIDIA, Meta, Google, OpenAI, Microsoft, Hugging Face, the Linux Foundation and others argued that open weights reduce barriers, increase competition, limit vendor lock-in and preserve customer control over data and deployment. The article therefore presents Huang's X post as public support for a coalition position, not as a manifesto he wrote alone.
Taiwan enterprises must include data sovereignty and procurement policy
On-premises inference does not automatically make a system secure. Verify weight sources and hashes, audit remote code and third-party quantizations, disable unnecessary telemetry, restrict outbound traffic and maintain a software bill of materials.
Taiwan's Ministry of Digital Affairs states that government agencies may not use DeepSeek services, apps or locally downloaded models. Private-sector use is not categorically prohibited, but government, critical-infrastructure and government-contracting teams should obtain written security and procurement approval before a PoC.
A practical adoption sequence
- Define data-location, source-country, license and hosting constraints.
- Use one or two hosted flagships to establish the quality ceiling.
- Compare Qwen3.8-27B, Nemotron Lightning, Muse, Gemma, gpt-oss, Ministral and TAIDE on the same local baseline.
- Improve RAG before fine-tuning when knowledge changes or citations matter.
- Load-test real context, output lengths and concurrency before purchasing hardware.
- Use a model gateway for routing, rollback, cost control, logs and security policy.
Enterprise selection FAQ for 2026
Which models belong in the first PoC?
Use one hosted flagship as the quality ceiling, then test Qwen3.8-27B for Chinese multimodality, Gemma 4 E4B or 26B and Muse Glimmer for local multimodality and supply-chain diversity, and TAIDE for Taiwan language. Add Nemotron Lightning when agent throughput matters.
What can realistically run on a 24GB or 32GB device?
Gemma 4 E2B and E4B are the clearest edge options; quantized Muse Glimmer explicitly targets 24GB/32GB devices, and TAIDE 12B is practical for Taiwan-language testing. Gemma 4 26B/31B and Qwen3.8-27B depend on quantization, context and hardware. A model that loads is not automatically a reliable multi-user service.
Which latest models actually provide downloadable weights?
Qwen 3.8, DeepSeek V4, Kimi K3, Nemotron, Muse Glimmer and Gemma 4 do, under different licenses and infrastructure requirements.
How should teams choose between DeepSeek Flash-0731 and Pro-0813?
Both are MIT-licensed with public weights; the difference is cost efficiency and deployment weight. Flash-0731 is 284B/13B-active and suits coding, tool calls and agent workloads; Pro-0813 is a trillion-scale MoE aimed at harder reasoning, and self-hosting it costs far more. Compare completion rate, retries and total task cost rather than headline token price.
When is API better than self-hosting?
API is usually efficient for low or variable volume and immediate frontier access. Self-hosting becomes attractive for strict data residency, predictable sustained utilization, fixed model versions or deep customization. Compare one-year TCO, not token price alone.
Can a Taiwan organization directly adopt a China-origin model?
Private-sector answers depend on data classification, regulation, contracts and supply-chain policy. Taiwan government restrictions are explicit for DeepSeek. Public-sector, critical-infrastructure and government-contracting teams should obtain written approval and keep alternative model routes ready.
Conclusion
The strongest strategy is to use hosted flagships as a quality ceiling, deployable open-weight models as the production baseline, and your own evaluation set as the decision authority. Qwen3.8-27B, Nemotron Lightning, Muse Glimmer, Gemma 4 and TAIDE are practical local candidates; DeepSeek V4, Qwen 2.4T, Kimi K3 and Nemotron Ultra are cluster projects; GLM 5.3 remains a service-first release. Governance, retrieval, permissions and monitoring matter as much as the model name.
Primary sources: Qwen3.8-27B, DeepSeek V4-Flash-0731, GLM 5.3, Kimi K3, Nemotron 3.5 Lightning, Muse Glimmer 30B, Gemma 4, and TAIDE.