Model News

JetBrains Mellum2.1: 12B MoE Open Model for Coding Agents

JetBrains released Mellum2.1, a 12B mixture-of-experts open model for coding agents trained on real repositories under Apache 2.0 license.

Headline card: JetBrains Mellum2.1: 12B MoE Open Model for Coding Agents
On this page
  1. What changed
  2. Why it matters
  3. What to test
  4. The conclusion

What changed

JetBrains released Mellum2.1, a 12B mixture-of-experts thinking model for coding agents. The model activates 2.5B parameters per token and ships under Apache 2.0 on Hugging Face. The architecture remains unchanged from Mellum2, released in June 2026. According to JetBrains, almost all improvements come from reinforcement learning conducted in real software environments rather than architectural changes.

The released checkpoint is Mellum2.1-12B-A2.5B-Thinking. It is a reasoning model that generates its chain of thought before answering. JetBrains targets three use cases: agent worker, general reasoning assistant, and private self-hosted deployment.

The model has 28 layers and 64 experts. A router activates 8 experts per token. Attention uses grouped-query attention with 32 query heads and 4 KV heads. Three of every four layers use a 1,024-token sliding window. Context length is 131,072 tokens, and the vocabulary has 98,304 tokens. Weights ship in bfloat16.

Why it matters

Mellum2.1 sits at a practical intersection: small enough to self-host on consumer-grade GPUs, yet trained specifically for agentic coding tasks. The model includes native support for tool use and repository exploration, which matters for developers building private AI systems that need to edit files and run tests without cloud dependencies.

The benchmark improvements signal where the model gained ground. SWE-bench Verified, a metric for software engineering agents, jumped from 2.0 to 47.0. This tracks the model’s ability to understand and fix real code issues. On LiveCodeBench v6, it scores 82.0, ahead of comparable open models like Qwen3.5-9B (75.4) and Gemma 4 E4B (69.4).

However, gaps remain. Qwen3.5-9B still leads on SWE-bench Pro (38.0 vs 28.0) and on pure knowledge benchmarks like GPQA Diamond (77.8 vs 64.6). This tells us that Mellum2.1 excels at code-specific reasoning but may struggle with broader domain knowledge or the hardest agentic tasks.

Speed matters for deployment. JetBrains says Mellum2.1 serves almost 2x the tokens of Qwen3.5-9B under heavy load on a single H200 GPU. For single requests with multi-token prediction (MTP), the speedup is about 1.6x. That efficiency comes from the mixture-of-experts design, which activates only 8 of 64 experts per token.

Locally, the model compresses down to 7.0 GB to 8.1 GB in GGUF format, making it practical for Ollama, llama.cpp, and LM Studio. The recommended quantization is Q4_K_M at 8.1 GB with 88% top-token match to the original. This changes the practical equation for teams that want coding agents without cloud services or vendor lock-in.

What to test

Before adopting Mellum2.1, check the following claims and gaps:

Benchmark fairness: All scores are self-reported by JetBrains using their own evaluation pipeline in thinking mode. Qwen and Gemma report different numbers on the same benchmarks when tested in their own pipelines. Test Mellum2.1 directly against Qwen3.5-9B on your own code samples and agentic workflows. The 47.0 on SWE-bench Verified is real progress, but real-world gains depend on your specific tasks.

Agentic capability limits: The model was trained with the Pi v0.73.1 harness using a 114K-token context window. Test whether it stays consistent across longer codebases or multi-file edits. The 28.0 on SWE-bench Pro (versus Qwen’s 38.0) suggests harder agentic tasks may still trip it up. Run your own agent sandbox to verify.

Speed claims: Throughput tests used an H200 under heavy load. Test actual latency and token throughput on your GPU stack. The multi-token prediction head is listed as coming soon for vLLM speculative decoding, so that 1.6x speedup is a target, not yet production-ready.

Safety regression: HarmBench improved from 21.5 to 8.5, lower is better. If safety matters for your use case, independently evaluate the model on your risk scenarios.

Quantization quality: GGUF builds range from 7.0 GB (MXFP4_MOE, 85.6% token match) to 12.9 GB (Q8_0, 96.1% token match). Test the 8.1 GB Q4_K_M quantization on your coding tasks before betting on it for production.

Real-world RL environment: The model was trained in real repositories with shell and file-editing tools across millions of sandboxes in thousands of environments. Test it on your actual tech stack. If your codebase uses unfamiliar patterns or older languages, the RL training may not transfer as well as benchmarks suggest.

The conclusion

Mellum2.1 represents a legitimate step forward for open coding models, particularly for agentic software engineering. The move from almost no SWE-bench Verified performance (2.0) to competitive scores (47.0) shows that RL in real environments works better than pure supervised fine-tuning. The efficiency gains from the 2.5B active parameter design make local deployment realistic.

But it is not the new winner across the board. Qwen3.5-9B still leads on the hardest agentic tasks and pure knowledge. Mellum2.1 excels on code generation and structured coding tasks. The right choice depends on your workload. If you need to run a coding agent locally without cloud dependencies, Mellum2.1 is worth testing. If you need broader reasoning or harder task solving, Qwen may still be safer.

Watch for the multi-token prediction head for vLLM and real-world adoption reports. Real teams using Mellum2.1 in production will surface gaps that benchmarks miss. The Apache 2.0 license removes vendor lock-in, so local experimentation costs almost nothing.

JetBrains is positioning this at the infrastructure layer of agentic systems, not as a replacement for larger closed models. That positioning is honest. It is fast, deployable, and trained for a specific job. Whether it handles your job is a question only your own testing can answer.

AI Tool Herald may earn a commission from some links on this site. It never changes what we report or recommend. Affiliate disclosure