Best Open-Weight AI Models 2026: Gemma 4 and Muse Glimmer
Gemma 4 and Muse Glimmer are the open-weight models worth self-hosting in 2026. Sizes, licences, published benchmarks, hosted costs, and an honest test for when self-hosting actually pays.
Published 23 September 2026 by Jake Hissitt, Founder of Stob.AI.
Open weights matter for three reasons: data that cannot leave your infrastructure, cost at extreme volume, and not being repriced by someone else's roadmap.
In 2026 the two families worth serious attention are Google's Gemma 4 and Meta's Muse Glimmer.
The shortlist Model Size Context Licence Cost Gemma 4 31B IT 30.7B dense 262,144 Apache 2.0 Free on the Gemini API, or self-host Gemma 4 26B-A4B IT 25.2B total / 3.8B active MoE 262,144 Apache 2.0 Free on the Gemini API, or self-host Gemma 4 12B Unified IT 11.95B 262,144 Apache 2.0 Host or self-host Gemma 4 E4B / E2B IT Edge sizes 131,072 Apache 2.0 Host or self-host Muse Glimmer 30B dense multimodal 131,072 Apache 2.0 Host or self-host; Meta cites Together AI at $0.35 / $1.50 Gemma 4: the strongest open reasoning line Google reports 85.2 MMLU-Pro, 89.2 AIME 2026, 84.3 GPQA Diamond and 80.0 LiveCodeBench v6 for Gemma 4 31B IT.
Those are serious numbers for a 30B model under Apache 2.0, and they put it ahead of much larger open models from a year ago.
The 26B-A4B mixture-of-experts variant reports 82.6 MMLU-Pro, 88.3 AIME 2026, 82.3 GPQA Diamond and 77.1 LiveCodeBench v6 while activating only 3.8B parameters per token — which is the interesting one if inference cost on your own hardware is the constraint.
One caution: the free Gemini API tier for Gemma 4 is genuinely free, but free capacity is not a production SLA.
If a workflow matters, self-host it or pay for hosted capacity.
Muse Glimmer: the local agent model Muse Glimmer is Meta's 30B dense multimodal open-weight release from 10 August 2026, aimed at always-on local agents: sequential tool use, private workflows and recovery from failed steps.
Meta cites Together AI at $0.35 in / $1.50 out per million tokens as one hosted example, but as an open-weight model the real cost depends on where you run it.
It is a different proposition from Muse Spark 1.3 , which is hosted and closed-weight.
Glimmer is what you reach for when the requirement is "this must run on our own hardware".