Meta’s Muse Glimmer Open‑Source Release Revives the Open‑Weight Llama Competition

Meta has unveiled Muse Glimmer, a 30‑billion‑parameter open‑source LLM under Apache 2.0, positioned for agent workloads and benchmarked against Google’s Gemma 4 and Alibaba’s Qwen, while highlighting hardware requirements, performance limits, and the broader strategic implications for U.S. AI policy.

21CTO
21CTO
21CTO
Meta’s Muse Glimmer Open‑Source Release Revives the Open‑Weight Llama Competition

Meta announced the release of Muse Glimmer, its first open‑source LLM in over a year, featuring 30 billion parameters and licensed under the permissive Apache 2.0 license. The model is derived from the larger, proprietary Muse Spark series and is aimed at local AI inference workloads such as agents, code assistants, and multimodal tools.

Benchmark Positioning

In Meta’s internal benchmarks, Muse Glimmer outperforms Google’s Gemma 4 in most cases and matches the performance of Alibaba’s Qwen models of comparable scale. However, the model remains far smaller than leading Chinese models such as Kimi K3, Qwen 3.8‑Max, or DeepSeek V4, limiting its ability to dominate the open‑weight arena.

Agent‑Centric Capabilities

The model’s training and evaluation focus on end‑to‑end task completion, reliable tool calling, multi‑step reasoning with failure recovery, multimodal input via an 1.8 billion‑parameter ViT‑G/14 encoder, and controllable inference intensity (low, medium, high, xhigh) through system prompts. Architecturally, it uses a 52‑layer dense Transformer with a “three‑region‑one‑global” attention pattern, a sliding window of 2048 tokens, GQA ratio 16:1, and a context length starting at 131 072 tokens, supporting over 100 languages. Knowledge is current up to 4 January 2026.

Hardware Requirements and Performance

At native BF16 precision, Muse Glimmer fits on a single Nvidia RTX Pro 6000 or AMD MI350P. After 4‑bit quantisation, the model shrinks from ~60 GB to just under 16 GB, allowing it to run on consumer GPUs with 20‑24 GB VRAM (e.g., RTX 30/4090, RX 7900 XT/XTX). Users with 16 GB cards must resort to 3‑bit quantisation due to memory limits.

Meta reports token generation speeds of 75–233 tok/s on GPUs with ~1.8 TB/s memory bandwidth (e.g., RTX 5090). On a MacBook Pro with an M5 Max CPU (≈614 GB/s bandwidth), speeds of 26.2–57.8 tok/s are expected, while typical laptops without high‑bandwidth memory may only achieve 6–14 tok/s. An Unsloth Studio test on a DGX Spark system measured ~12.2 tok/s, though DSpark support is not yet available.

Availability

The weights are hosted on Hugging Face and can be downloaded directly or via local inference platforms such as Ollama and LM Studio.

Strategic Context

Meta’s CEO Mark Zuckerberg warned that maintaining U.S. leadership in open‑source AI requires policy shifts around data usage and open‑software promotion, criticizing the concentration of AI power in a few firms. He emphasized that superintelligence should be widely distributed rather than centralized, envisioning personal AI agents that empower individuals.

Future plans include an upcoming open‑weight version of Muse Spark 1.2, which Meta claims will be its most competitive model against OpenAI, Anthropic, and Google, though it still trails Chinese models in raw performance.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

LLMopen sourceBenchmarkMetahardware requirementsagent modelsMuse Glimmer
21CTO
Written by

21CTO

21CTO (21CTO.com) offers developers community, training, and services, making it your go‑to learning and service platform.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.