Skip to content

MiniMax · 428B · transformer-moe

2026-06-111.0M context428B params

Use Cases

chatcodereasoningmultilingualvisiontools

Quantization Options

QuantBitsVRAMQualityStatus
Q2_Krec2163.0 GBModerate
Q4_K_M4284.0 GBGood
Q8_08473.0 GBExcellent

About this model

MiniMax M3 is the first open-weight model to combine frontier-grade coding, a 1M-token context window, and native multimodality — image and video understanding are built into the architecture, not bolted on. A 428B parameter Mixture-of-Experts with ~23B active per token, it uses MiniMax Sparse Attention to cut per-token compute at long context to roughly 1/20 of the prior generation. Company-reported scores: 59.0 on SWE-bench Pro and 80.5 on SWE-bench Verified. On Ollama it ships only as the `minimax-m3:cloud` tag (512K context); local runs use community GGUFs. Thanks to the low active-parameter count it's the most attainable of the 2026 frontier MoEs: the ~143 GB 2-bit quant fits a 192 GB Mac Studio, and token throughput on unified memory is good for the size class. Note the license requires attribution and separate authorization for companies above $20M revenue.

Benchmarks

59.0
swe-bench-pro
80.5
swe-bench-verified