Skip to content
Original test by AI Radar model test

Meta Muse Glimmer 30B running locally on a MacBook Pro with Apple M4 Max

Author AI Radar

I tried Muse Glimmer 30B locally without a cloud API. I tested Czech, working with an image, structured JSON, a logic problem and drafting Python code.

Introduction

Muse Glimmer 30B is a multimodal local model that can run without sending text and images to the cloud. In this test, it ran in LM Studio on a MacBook Pro with an Apple M4 Max chip and 36 GB unified memory.

I focused on practical use: Czech, precise instructions, a logic problem, image analysis, generating Python code and structured JSON. The model is usable on this configuration, but it is slower on more complex tasks.

Test hardware

The model fits on the computer, but 36 GB of memory does not leave much headroom for other demanding applications. I used a context of 8192 tokens.

The exact quantization type was not listed in LM Studio, so I do not label it as, for example, Q4_K_M without further verification.

Computer
MacBook Pro 14"
Processor
Apple M4 Max
Memory
36 GB unified memory
Operating system
macOS Tahoe 26.6.1
Model
Muse Glimmer 30B
Format
GGUF, K-quant
Size on disk
18.16 GB
Test context
8192 tokens

Security configuration

The server listened only on 127.0.0.1, local network sharing was disabled and authentication was not enabled because the server was not accessible from outside the computer.

I also disabled CORS, MCP connections and automatic model loading, and limited concurrent predictions to one. After the test, I stopped the server and unloaded the model from memory.

What I verified

The GET /v1/models endpoint confirmed that LM Studio was serving the loaded model meta/muse-glimmer locally through an OpenAI-compatible API. I did not independently make a separate POST request to the chat API.

  • Czech response: the model answered correctly and concisely.
  • Exact format: it provided exactly three Czech bullet points with a limit of 12 words.
  • Logic problem: it correctly solved the problem involving three switches and a light bulb.
  • Image: it correctly described a screenshot of the LM Studio interface.
  • Python: it drafted a function for normalizing a Czech phone number, but I did not run the code.
  • JSON: it returned a syntactically valid object with the required fields.

Speed

The biggest practical difference compared with cloud models appeared in reasoning. Simpler responses were usable, but drafting the Python code took several minutes.

Czech response
approximately 4.51 tokens/s
Logic problem
approximately 7.22 tokens/s
Image analysis
approximately 5.51 tokens/s
Structured JSON
approximately 10.83 tokens/s

Limitations

No standardized benchmark, long PDF test, RAG, work with a large database or actual tool calls were performed. I also did not conduct a direct comparison with other local models.

The generated Python code was not run in an interpreter. The result therefore confirms the ability to draft code, not that it works without errors.

AI Radar verdict

What the test showed

Muse Glimmer 30B can be run locally in LM Studio for practical use on a MacBook Pro with M4 Max and 36 GB of memory. In the test, it responded correctly in Czech, followed precise instructions, solved a logic problem, analyzed a screenshot and produced valid JSON. The biggest limitation is slower reasoning on complex requests.