Benchmarks Architecture Text Decoder Perception Encoder Transformers Text-only Inference Prompting the model with images and text Video Inference Multimodal tool calling Object Detection Llama.cpp Speculative Decoding Speculative Decoding with transformers Speculative Decoding with llama.cpp Inference Endpoints Support for Muse Glimmer vLLM with transformers backend Fine-tuning with TRL Demos Connect OpenClaw to Muse Glimmer Hey Muse Glimmer, quantize yourself Hey Muse Glimmer, deploy yourself Hey Muse Glimmer, optimize yourself Hey Muse Glimmer, research the Hub Wrapping Up Great news from the OGs of open source LLMs! Muse Glimmer, released today, is Meta’s new multimodal model, especially designed for local agentic use cases. Distilled from Muse to 30B parameters, and released under the Apache 2.0 license, it’s ideal deploying locally for privacy, reducing costs, or just hacking around. It’s intended for privacy-aware applications such as coding, document analysis, personal assistants, Claw- or Hermes-like setups.
Meta is back with Muse Glimmer: local, agentic, multimodal, and open source
We’re on a journey to advance and democratize artificial intelligence through open source and open science.

Key takeaways
Meta has released Muse Glimmer, a multimodal model for local, agentic use cases, with a 30B parameter count and Apache 2.0 license.
- Muse Glimmer is designed for local use, reducing costs and improving privacy.
- It uses a 2B ViT-like image encoder and a transformer-based architecture.
- The model supports video inference, multimodal tool calling, and object detection.
- DFlash provides a drafter for faster generation with some memory cost.
Summarised automatically by AI from the original article by Hugging Face Blog. AI can make mistakes, so check the original for details.
Story details
- Published
- By
- Pedro Cuenca, merve, ben burtenshaw
- Format
- Article
- Original
- huggingface.co ↗
Related reporting
2 stories on this topic from 2 sources, oldest first.
- Hugging Face Blog Meta is back with Muse Glimmer: local, agentic, multimodal, and open source (this story)



