Runtime as featured inForbesRead the article

The best open source AI models, ranked and compared

A guide to the top open-weight language models for AI agents. Compare parameters, context, license, and input type, then run any of them inside your own cloud with Runtime.

Updated · 11 models

The best open source LLMs for agents

Every model here publishes its weights. We rank them for agent work: tool use over long tasks, context length, license, and what it costs to run at volume.

1Z.ai logo
Best overall

GLM-5.3 keeps the GLM-5.2 base model and puts every gain into post-training. Z.ai reports open-source state of the art on Terminal Bench 3.0 and Agents' Last Exam, and a 50% gain over GLM-5.2 on its in-house code benchmark.

  • 1M-token context
  • Top open-weight scores on agentic benchmarks
  • Built for long, multi-step tool use

Good for

Payment reconciliation: Top open-weight agentic scores, so it reliably chains ledger queries, processor APIs, and file parsing across a long run.

Z.ai's strongest open-weight model, post-trained for complex coding and long-horizon agent work.

Context1M tokens
Parameters753B
Provider
Z.ai
Architecture
Mixture-of-experts
Parameters
753B total
Context
1M tokens | 786K words
License
GLM-5.3 License
Input
Text
Released
Aug 2026

Harnesses in Runtime

Claude CodeOpenCodeCline

Provider-reported benchmarks

Terminal Bench 2.1 88.2Toolathlon Verified 73.0CyberGym 84.5
2Moonshot AI logo

Kimi K3

Moonshot AI

Most capable

Kimi K3 is a 2.8T-parameter mixture-of-experts model that activates 16 of 896 experts per token. Moonshot built it for long-horizon coding and agentic knowledge work, and it reads text, images, and video in the same model.

  • 2.8T parameters
  • Text, image, and video input
  • 1M-token context

Good for

Partner-bank RFIs: Native text, image, and video input with a 1M-token context, so it reads the whole RFI packet in one pass.

The first open 3T-class model, natively multimodal with a 1M-token context.

Context1M tokens
Parameters2.8T
Provider
Moonshot AI
Architecture
Mixture-of-experts
Parameters
2.8T total
Context
1M tokens | 786K words
License
Kimi K3 License
Input
Text, Image, Video
Released
Jun 2026

Harnesses in Runtime

Claude CodeOpenCodeCline

Provider-reported benchmarks

Terminal Bench 2.1 88.3Toolathlon Verified 76.5
3DeepSeek logo
Best for reasoning MIT

DeepSeek-V4-Pro-0813 supersedes the V4 Pro preview with much better agentic performance in production settings. It adds a DSpark speculative decoding module and is released under MIT, so commercial use and fine-tuning are straightforward.

  • MIT license
  • 1M-token context
  • Strong reasoning and tool use

Good for

Payment reconciliation: Strong reasoning, a 1M-token context for whole settlement files, and an MIT license you can self-host without legal review.

The official DeepSeek V4 Pro release, with stronger agentic skills and an MIT license.

Context1M tokens
Parameters1.6T
Provider
DeepSeek
Architecture
Mixture-of-experts
Parameters
1.6T total
Context
1M tokens | 786K words
License
MIT Commercial use
Input
Text
Released
Aug 2026

Harnesses in Runtime

Claude CodeOpenCodeCline

Provider-reported benchmarks

Terminal Bench 2.1 87.9Toolathlon Verified 74.1CyberGym 83.3
4Z.ai logo
Best value MIT

GLM-5.3 Flash starts from a new base model with hybrid sparse and linear attention, which cuts long-context serving costs. With 18B active parameters it is cheap to run, and Z.ai reports it approaching Claude Opus 4.8 on coding and agentic benchmarks.

  • 18B active parameters
  • MIT license
  • Image input and 1M context

Good for

Partner-bank RFIs: Image input at 18B active parameters and an MIT license, a cost-effective default for high RFI volume.

A natively multimodal GLM that Z.ai says beats GLM-5.2 at one-tenth the price.

Context1M tokens
Parameters320B
Active18B
Provider
Z.ai
Architecture
Mixture-of-experts
Parameters
320B total · 18B active
Context
1M tokens | 786K words
License
MIT Commercial use
Input
Text, Image
Released
Aug 2026

Harnesses in Runtime

Claude CodeOpenCodeCline
5DeepSeek logo
Fastest MIT

DeepSeek V4.1 Flash uses a causal encoder-decoder design that lets the decoder reuse a compressed KV cache, about 4x smaller than V4 Flash. It reads images and text, supports a 1M-token context, and lets you dial reasoning effort up or down.

  • 8B active parameters
  • About 4x smaller KV cache than V4 Flash
  • MIT license

Good for

Transaction monitoring: Only 8B active parameters and a much smaller KV cache, so it triages thousands of alerts quickly and cheaply.

A multimodal MoE that activates only 8B parameters per token and shrinks the KV cache about 4x.

Context1M tokens
Parameters552B
Active8B
Provider
DeepSeek
Architecture
Mixture-of-experts
Parameters
552B total · 8B active
Context
1M tokens | 786K words
License
MIT Commercial use
Input
Text, Image
Released
Sep 2026

Harnesses in Runtime

Claude CodeOpenCodeCline
6Moonshot AI logo

Kimi K2.7 Code

Moonshot AI

Best for coding

Kimi K2.7 Code builds on Kimi K2.6 with large gains on real-world, multi-step coding tasks. Moonshot reports about 30% fewer thinking tokens than K2.6, which lowers cost on long runs.

  • Built for long-horizon coding
  • About 30% fewer thinking tokens than K2.6
  • 256K context

A coding-focused Kimi tuned for long-horizon software engineering.

Context256K tokens
Parameters1T
Active32B
Provider
Moonshot AI
Architecture
Mixture-of-experts
Parameters
1T total · 32B active
Context
256K tokens | 197K words
License
Modified MIT
Input
Text, Image
Released
Jun 2026

Harnesses in Runtime

Claude CodeOpenCodeCline
7Qwen (Alibaba) logo

Qwen3.8 2.4T-A95B

Qwen (Alibaba)

Best for multilingual

Qwen3.8 2.4T-A95B is the open-weight model behind Qwen3.8-Max. Alibaba reports large gains in coding, professional work, and long-horizon agent tasks over Qwen3.5 and 3.6, with tunable reasoning effort.

  • 2.4T total, 95B active
  • Qwen-Max class, open weights
  • Tunable reasoning effort

Good for

AML and compliance investigations: Qwen-Max-class quality with broad multilingual coverage for cross-border alerts.

The first Qwen-Max-class model released with open weights.

Context256K tokens
Parameters2.4T
Active95B
Provider
Qwen (Alibaba)
Architecture
Mixture-of-experts
Parameters
2.4T total · 95B active
Context
256K tokens | 197K words
License
Qwen3.8-Max License
Input
Text
Released
Aug 2026

Harnesses in Runtime

OpenCodeCline
8MiniMax logo

MiniMax M3

MiniMax

Best multimodal

MiniMax M3 trains on text, images, and video from the first step. Its MiniMax Sparse Attention gives 9x faster prefill and 15x faster decode than M2 at 1M context, and it has enabled, adaptive, and disabled reasoning modes.

  • Text, image, and video input
  • 9x prefill and 15x decode speedup vs M2 at 1M context
  • Adaptive reasoning mode

Good for

Partner-bank RFIs: Multimodal with sparse attention built for million-token contexts, useful for long correspondence threads.

A natively multimodal model with sparse attention built for million-token contexts.

Context1M tokens
Parameters428B
Active23B
Provider
MiniMax
Architecture
Mixture-of-experts
Parameters
428B total · 23B active
Context
1M tokens | 786K words
License
MiniMax Community License
Input
Text, Image, Video
Released
Jun 2026

Harnesses in Runtime

Claude CodeOpenCodeCline
9Qwen (Alibaba) logo

Qwen3.8 27B

Qwen (Alibaba)

Best to self-host Apache 2.0

Qwen3.8 27B brings the Qwen3.8 generation to a deployment-friendly dense model. It understands images and video, thinking can be switched off per request, and the Apache 2.0 license makes it simple to run in your own cloud.

  • Apache 2.0 license
  • Small enough to self-host
  • Image and video input

Good for

Payment reconciliation: Small enough to run on your own GPUs under Apache 2.0 when ledger data cannot leave your VPC.

A compact dense vision-language model under Apache 2.0.

Context256K tokens
Parameters27B
Provider
Qwen (Alibaba)
Architecture
Dense
Parameters
27B total
Context
256K tokens | 197K words
License
Apache 2.0 Commercial use
Input
Text, Image, Video
Released
Aug 2026

Harnesses in Runtime

OpenCodeCline
10Qwen (Alibaba) logo

Qwen3.8 Flash Next

Qwen (Alibaba)

Most efficient long context

Qwen3.8 Flash Next pairs Gated DeltaNet with Qwen Sparse Attention to cut long-context latency, and adds n-gram embeddings for cheaper parameter scaling. Alibaba positions it as a preview of the Qwen4 architecture.

  • Hybrid attention for low long-context latency
  • Image input
  • Preview of the Qwen4 architecture

An experimental preview of the architecture that will underpin Qwen4.

Context256K tokens
Parameters180B
Provider
Qwen (Alibaba)
Architecture
Mixture-of-experts
Parameters
180B total
Context
256K tokens | 197K words
License
Qwen Community License
Input
Text, Image
Released
Aug 2026

Harnesses in Runtime

OpenCodeCline
11OpenAI logo
Best on a single GPU Apache 2.0

gpt-oss-120b is OpenAI's open-weight reasoning model under Apache 2.0. It activates 5.1B parameters per token and fits on a single 80GB GPU, which makes it a common choice for private, low-cost deployments.

  • Apache 2.0 license
  • Runs on one 80GB GPU
  • Configurable reasoning effort

Good for

Transaction monitoring: Apache 2.0 and fits on a single 80GB GPU, a simple way to run triage entirely inside your network.

OpenAI's open-weight reasoning model, small enough to run on a single 80GB GPU.

Context128K tokens
Parameters117B
Active5.1B
Provider
OpenAI
Architecture
Mixture-of-experts
Parameters
117B total · 5.1B active
Context
128K tokens | 98K words
License
Apache 2.0 Commercial use
Input
Text
Released
Aug 2025

Harnesses in Runtime

CodexOpenCodeCline

Benchmark scores are reported by each model provider on its Hugging Face model card. They are not independently verified by Runtime.

Open source models compared

Specs side by side. Filter by license or input type and sort by context or release date.

Filter:11 of 11 models
ParametersLicenseInputGood for
1Z.ai logoGLM-5.3Best overall753B1M tokensGLM-5.3 LicensetextAug 2026Payment reconciliationChargebacks and disputesMerchant underwriting
2Moonshot AI logoKimi K3Most capable2.8T1M tokensKimi K3 Licensetext, image, videoJun 2026Partner-bank RFIsChargebacks and disputesMerchant underwriting
3DeepSeek logoDeepSeek V4 ProBest for reasoning1.6T1M tokensMITtextAug 2026Payment reconciliationMerchant underwritingAML and compliance investigations
4Z.ai logoGLM-5.3 FlashBest value320B / 18B1M tokensMITtext, imageAug 2026Partner-bank RFIsTransaction monitoringKYC and KYB checks
5DeepSeek logoDeepSeek V4.1 FlashFastest552B / 8B1M tokensMITtext, imageSep 2026Transaction monitoringKYC and KYB checksPayments customer support
6Moonshot AI logoKimi K2.7 CodeBest for coding1T / 32B256K tokensModified MITtext, imageJun 2026
7Qwen (Alibaba) logoQwen3.8 2.4T-A95BBest for multilingual2.4T / 95B256K tokensQwen3.8-Max LicensetextAug 2026AML and compliance investigationsPayments customer support
8MiniMax logoMiniMax M3Best multimodal428B / 23B1M tokensMiniMax Community Licensetext, image, videoJun 2026Partner-bank RFIsChargebacks and disputes
9Qwen (Alibaba) logoQwen3.8 27BBest to self-host27B256K tokensApache 2.0text, image, videoAug 2026Payment reconciliationKYC and KYB checks
10Qwen (Alibaba) logoQwen3.8 Flash NextMost efficient long context180B256K tokensQwen Community Licensetext, imageAug 2026
11OpenAI logogpt-oss-120bBest on a single GPU117B / 5.1B128K tokensApache 2.0textAug 2025Transaction monitoring

Parameters show total / active per token for mixture-of-experts models.

Open source vs frontier models

Frontier models still lead on the hardest tasks. Open-weight models win on cost, control, and customization, and the gap narrows with every release.

DimensionOpen source modelsFrontier models
CostLow per-token cost, often a fraction of frontier pricing.Premium pricing on the strongest tiers.
Data controlRun in your own cloud, so card data and PII never leave your environment.Hosted by the provider under their data policy.
CustomizationFine-tune the weights on your own cases and runbooks.Limited to prompting and provider tuning.
Peak capabilityClose behind on most agentic benchmarks, and ahead on cost per task.Still leads on the hardest tasks.
HostingSelf-host, use Bedrock or Vertex AI, or a hosting provider.Provider API only.

Run open models inside your own cloud

Runtime is model-neutral. Pick an open-weight model per agent or per skill, keep sensitive data in your environment, and let evals tell you when a cheaper model is good enough.

Your inference, your cloud

Serve models from Amazon Bedrock, Google Vertex AI, or your own vLLM or SGLang endpoint. Card data and PII stay in your environment.

Any harness

Run open models through OpenCode or Cline, or through Claude Code with providers that offer an Anthropic-compatible API, like Z.ai, Moonshot, DeepSeek, and MiniMax.

Routed by evals

Test each skill against real cases, then move routine work to the smallest model that still passes.

Understanding open source AI models

What are open source AI models?

Open source AI models are large language models whose weights are published under a license that lets you download, run, and usually modify them. Instead of calling one provider's API, you can host the model yourself, use a managed service like Bedrock or Vertex AI, or run it inside a platform like Runtime.

Open weights vs fully open source

Most models called open source are open weight: the trained weights are public, but the training data and recipe may not be. MIT and Apache 2.0 are permissive and safe for commercial work. Custom community licenses can add conditions, so read them before you self-host or fine-tune.

How mixture-of-experts keeps costs down

Most of the strongest open models use a mixture-of-experts design, which sends each token through a small slice of a much larger network. That is why a model can have hundreds of billions of parameters but only activate 8B to 95B per token, keeping quality high and inference cheap.

How to choose an open source model

Start with the task. Match the context window to your inputs, check the license if you plan to self-host, and weigh cost against the quality the work needs. For regulated data, favor a model you can run in your own cloud, then test a few against your real cases before you commit.

Open source AI models FAQ

What is the best open source AI model right now?

+

As of October 1, 2026, our pick for agent work is GLM-5.3 from Z.ai. Z.ai's strongest open-weight model, post-trained for complex coding and long-horizon agent work. It has a 1M-token context window and is released under the GLM-5.3 License.

Which open source model is best for coding?

+

Kimi K2.7 Code from Moonshot AI is the coding-focused pick in our ranking. A coding-focused Kimi tuned for long-horizon software engineering. GLM-5.3 is also a strong choice for long agentic coding tasks.

Which open source models can I use commercially?

+

Models under MIT or Apache 2.0 are permissive and safe for commercial use and fine-tuning. In this list that includes DeepSeek V4 Pro (MIT), GLM-5.3 Flash (MIT), DeepSeek V4.1 Flash (MIT), Qwen3.8 27B (Apache 2.0), gpt-oss-120b (Apache 2.0). Other models use custom community licenses, so read the terms before you self-host or fine-tune.

Are open source models as good as frontier models?

+

On the hardest tasks, frontier models from Anthropic, OpenAI, and Google still lead. The gap is now small on many agentic benchmarks, and open-weight models win on cost, privacy, and control because you can run them in your own cloud.

What is the difference between open source and open weight?

+

Open-weight models publish the trained weights so you can run and fine-tune them, but may not release the training data or code. Fully open source models publish those too. Most models called open source today are open weight, so the license is what matters in practice.

Which open source model is best to self-host?

+

Qwen3.8 27B is our pick for self-hosting: a compact dense vision-language model under Apache 2.0. For a single-GPU setup, gpt-oss-120b fits on one 80GB GPU.

Can I run these open source models in Runtime?

+

Yes. Runtime is model-neutral. Agents can use open-weight models served from Amazon Bedrock, Google Vertex AI, or your own inference endpoint in your cloud, through harnesses like OpenCode. Evals check each skill against real cases so you can move routine work to cheaper models.

Run open source models in Runtime

Pick a model, connect your tools, and put an agent to work inside your own cloud.