The frontier models have led the pack for a while now. It seems like the big players of Anthropic and OpenAI keep leapfrogging each other by a couple points in benchmark scores every other month. But, a trend we are starting to see is that open weight models are improving by leaps and bounds. They don’t hold the lead and probably won’t for a while, but the fact that open models are scaring the leaders is something to think about.
The Current Landscape
The featured image at the top of this page shows the current state of things according to artificialanalysis.ai. I have removed some of the other popular models like Grok and Gemini to showcase where the best open weight models sit compared to the frontier. Some things that should probably jump out to you are GLM 5.2 and Deepseek V4 Flash holding their own against the frontier models of GPT 5.6 Luna and Claude Sonnet 5. Another thing to note is that Minimax M3 and Mimo 2.5 Pro are somewhat older models by comparison to the leaders but they are keeping up enough to be useful. The biggest surprise is seeing Kimi K3 sitting at 4th place with only a few points separating it from Fable and Sol. But, let’s look at some of the models specifically in relation to their older iterations.
Minimax
Minimax M3 was one of my favorite workhorses for quite a while. It may not have scored as high in coding when compared to GLM 5.2, but one important thing to recognize is that it supports image inputs and GLM doesn’t. This is especially useful when doing frontend work where the model needs to visually inspect the changes it made using a browser automation like Playwright or agent-browser. The other thing to notice is that it’s a pretty significant jump from its M2.7 predecessor that was released just 2 months prior.
GLM
The GLM models have been popular for ages. There was even speculation that GLM 4.7 was powering OpenCode’s “Big Pickle” model for a while, though I’m personally thinking that Big Pickle is just a facade for a model that OpenCode chooses and changes as time goes on. While GLM 5.2 doesn’t have image input capabilities, it’s a very capable model when it comes to coding. The interesting note is that GLM jumped in 11 points of intelligence from 5.1 to 5.2 and that was only over the course of 2 months as well.
As you can see in the video above, I personally have seen it go off the rails before and loop on some nonsense but almost all the models I have tried do that from time to time, even the frontier ones. It’s just less often with the frontier models.
Kimi
The model taking the internet by storm is Kimi K3. The fact that it’s outscoring Opus 4.8 and GPT 5.6 Terra is amazing. The intelligence comes with a cost though. It is fairly expensive on services like OpenRouter and it likes to burn tokens while also being somewhat slower than comparable models. The model is so costly and popular that their own subscription plans over on the Kimi Code website are sold out and waitlisted.
Another example is Ollama, a site with a subscription plan that allows you to choose from a significant amount of open models. They don’t even allow you to use Kimi K3 without paying for extra usage credits. The best place I have found for trying out Kimi K3 is OpenCode Go, but they bill Kimi K3 at 2x token usage. So, it’s pretty easy to use up that usage within a week or two. However, it seems to be worth it and continues the theme of rapid improvement. They made their 13 point intelligence jump within 3 months of their previous Kimi K2.6 release.
Deepseek
Deepseek is a very interesting case. They were one of the first open weight models to begin really competing with the frontier. However, they caught a lot of flak for distilling the frontier models to do so. I think that is a bit hypocritical of the frontier model providers to take offense to such a thing, though. After all, the frontier models ingested the entire internet and trained their models on it with seemingly zero regard for copyright. If it’s ok for them to do that, I’m not sure I see a problem with people distilling models from information the frontier providers don’t truly own. But, that is content for another rant.
The thing I really want to shout out is that Deepseek is getting its latest improvement from fine tuning their existing Flash model rather than training up an entirely new model. The V4 family of Flash and Pro were already significant jumps from their previous V3.2. But, the newest version of V4 Flash 0731 has massively leapfrogged their own V4 Pro model. It may not be the most intelligent model in the open weight category, but if they were able to fine tune Flash in such a way, a fine tuning of their Pro model will be a serious contender.
The Future Is Bright
Just as it seems like the frontier models are hitting diminishing returns, the open weight models are catching up. This is good news for everyone besides the frontier providers. It even seems like the frontier providers are getting scared and trying to create some kind of regulatory capture to prevent competitors from rising up. It is in everyone’s best interest to make sure the landscape stays competitive and not concentrated in the hands of a few early winners who want to pull the ladder up behind them.
GitKraken MCP
GitKraken Insights