Weights first
“Live” means a downloadable checkpoint exists. API access alone does not qualify.
REVIEWED · JUL 20, 2026
OPEN-WEIGHT INTELLIGENCE · EVIDENCE FIRST
Open-weight models, compared without the hand-waving. See what is downloadable, what the license really permits, and what it takes to self-host.
models tracked
weights live
openness tiers
largest context
Provider benchmarks use different harnesses and are shown as evidence notes, never as a cross-model ranking.
| Compare | Model | Status | Scale | Context | Inputs | Reasoning + agents | Openness | Self-hosting | Source |
|---|---|---|---|---|---|---|---|---|---|
| Kimi K3Moonshot AI · Jul 16, 2026 | Weights Jul 27 | 2.8TUndisclosed active | 1MAdvertised | Max at launch; lower efforts plannedCoding, tools, Agent Swarm | License pendingTBD with weights | API today; weights promised Jul 27 | Official ↗ | ||
| InklingThinking Machines · Jul 15, 2026 | Weights live | 975B41B active | 1M64K / 256K on Tinker | Controllable effortCoding, tools, fine-tuning on Tinker | Use-restrictedApache 2.0 + Model AUP | ≥600 GB quantized; ≥2 TB BF16 | Official ↗ | ||
| GLM-5.2Z.ai · Jun 16, 2026 | Weights live | 744B40B active | 1MNative long-horizon target | Off / high / maxTool calling, long-horizon engineering | PermissiveMIT | Cluster class; BF16 and FP8 | Official ↗ | ||
| Kimi K2.7 CodeMoonshot AI · Jun 12, 2026 | Weights live | 1T32B active | 256KNative | Forced, preserved across turnsLong-horizon coding and tools | Use-restrictedModified MIT | Large multi-GPU; native INT4 | Official ↗ | ||
| Nemotron 3 UltraNVIDIA · Jun 4, 2026 | Weights live | 550B55B active | 1MRULER tested | Off / regular / mediumStructured tools, delegation, recovery | PermissiveOpenMDW 1.1 | 8×H200/B200 or 16×H100 BF16 | Official ↗ | ||
| Gemma 4 12BGoogle DeepMind · Jun 3, 2026 | Weights live | 12B12B dense active | 256KNative | Configurable thinkingNative function calling, structured JSON | PermissiveApache 2.0 | Laptop class; targets 16 GB memory | Official ↗ | ||
| MiniMax M3MiniMax · Jun 1, 2026 | Weights live | ~428B~23B active | 1MNative | Off / enabled / adaptiveCoding, computer use, long-horizon agents | Use-restrictedMiniMax Community | Cluster class | Official ↗ | ||
| Step 3.7 FlashStepFun · May 29, 2026 | Weights live | 196B + 1.8B vision~11B active | 256KNative | Low / medium / highVisual search, tools, GUI workflows | PermissiveApache 2.0 | ~120 GB minimum unified memory | Official ↗ | ||
| Mistral Medium 3.5Mistral AI · Apr 28, 2026 | Weights live | 128B128B dense active | 256KNative | Configurable effortFunction calling, JSON, coding | Use-restrictedModified MIT | As few as 4 GPUs (vendor) | Official ↗ | ||
| DeepSeek V4 ProDeepSeek · Apr 24, 2026 | Preview · weights live | 1.6T49B active | 1MThink Max recommends ≥384K | Off / think / think maxCoding, reasoning, tools | PermissiveMIT | Cluster class; mixed FP4 / FP8 | Official ↗ | ||
| Qwen3.6 27BAlibaba Qwen · Apr 22, 2026 | Weights live | 27B27B dense active | 256KExtensible to ~1M | Thinking / non-thinkingTools, repository reasoning, coding | PermissiveApache 2.0 | Workstation/server; 8 GPUs for full 256K | Official ↗ | ||
| GLM-4.7-FlashZ.ai · Jan 19, 2026 | Weights live | 30B3B active | ~200K202,752 positions | Reasoning with preserved thinkingTools and coding agents | PermissiveMIT | Local/workstation friendly | Official ↗ | ||
| gpt-oss-120bOpenAI · Aug 5, 2025 | Weights live | 117B5.1B active | 128KNative | Low / medium / highTools, functions, structured output | PermissiveApache 2.0 + usage policy | Single 80 GB GPU | Official ↗ | ||
| Llama 4 MaverickMeta · Apr 5, 2025 | Weights live | 400B17B active | 1MNative claim | General reasoningTool support depends on runner | Custom licenseLlama 4 Community | DGX-class FP8 deployment | Official ↗ |
COMPARE · 2/2
Inkling ↔ GLM-5.2
COMPARISON · 02
METHODOLOGY · 03
“Live” means a downloadable checkpoint exists. API access alone does not qualify.
Permissive, use-restricted, custom, and pending terms are kept separate.
Provider benchmarks remain labeled and are not normalized into a misleading leaderboard.
A model is only “runnable” in context: laptop, workstation, or cluster-class.