Meta Muse Glimmer 30B running locally on a MacBook Pro with Apple M4 Max
I tried Muse Glimmer 30B locally without a cloud API. I tested Czech, working with an image, structured JSON, a logic problem and drafting Python code.
Introduction
Muse Glimmer 30B is a multimodal local model that can run without sending text and images to the cloud. In this test, it ran in LM Studio on a MacBook Pro with an Apple M4 Max chip and 36 GB unified memory.
I focused on practical use: Czech, precise instructions, a logic problem, image analysis, generating Python code and structured JSON. The model is usable on this configuration, but it is slower on more complex tasks.
Test hardware
The model fits on the computer, but 36 GB of memory does not leave much headroom for other demanding applications. I used a context of 8192 tokens.
The exact quantization type was not listed in LM Studio, so I do not label it as, for example, Q4_K_M without further verification.
- Computer
- MacBook Pro 14"
- Processor
- Apple M4 Max
- Memory
- 36 GB unified memory
- Operating system
- macOS Tahoe 26.6.1
- Model
- Muse Glimmer 30B
- Format
- GGUF, K-quant
- Size on disk
- 18.16 GB
- Test context
- 8192 tokens
Security configuration
The server listened only on 127.0.0.1, local network sharing was disabled and authentication was not enabled because the server was not accessible from outside the computer.
I also disabled CORS, MCP connections and automatic model loading, and limited concurrent predictions to one. After the test, I stopped the server and unloaded the model from memory.
What I verified
The GET /v1/models endpoint confirmed that LM Studio was serving the loaded model meta/muse-glimmer locally through an OpenAI-compatible API. I did not independently make a separate POST request to the chat API.
- Czech response: the model answered correctly and concisely.
- Exact format: it provided exactly three Czech bullet points with a limit of 12 words.
- Logic problem: it correctly solved the problem involving three switches and a light bulb.
- Image: it correctly described a screenshot of the LM Studio interface.
- Python: it drafted a function for normalizing a Czech phone number, but I did not run the code.
- JSON: it returned a syntactically valid object with the required fields.
Speed
The biggest practical difference compared with cloud models appeared in reasoning. Simpler responses were usable, but drafting the Python code took several minutes.
- Czech response
- approximately 4.51 tokens/s
- Logic problem
- approximately 7.22 tokens/s
- Image analysis
- approximately 5.51 tokens/s
- Structured JSON
- approximately 10.83 tokens/s
Limitations
No standardized benchmark, long PDF test, RAG, work with a large database or actual tool calls were performed. I also did not conduct a direct comparison with other local models.
The generated Python code was not run in an interpreter. The result therefore confirms the ability to draft code, not that it works without errors.
AI Radar verdict
What the test showed
Muse Glimmer 30B can be run locally in LM Studio for practical use on a MacBook Pro with M4 Max and 36 GB of memory. In the test, it responded correctly in Czech, followed precise instructions, solved a logic problem, analyzed a screenshot and produced valid JSON. The biggest limitation is slower reasoning on complex requests.