Agentic Research

oMLX Installation and Usage Guide: Running MLX Models Natively on Apple Silicon

2026/05/1011 min readBryan Chan閱讀中文原文
TopicsoMLXMLXApple SiliconLocal LLMMac

What is oMLX?

oMLX is a desktop application based on the Apple MLX framework that lets you natively run MLX-format LLMs on Apple Silicon Macs. MLX is Apple's official machine learning framework, optimized specifically for M-series chips.

  • Native Apple Silicon: Leverages the M3 Ultra's GPU + Neural Engine
  • Unified Memory: Models load directly into 512GB unified memory
  • No Rosetta required: Native ARM64, faster than llama.cpp

Installation

# Download the Latest Version
# https://github.com/open-mlx/oMLX/releases

# Or via Homebrew (if there is a cask)
brew install --cask omlx

After downloading the .dmg, drag it into Applications.


Model Import

oMLX supports .safetensors and MLX format models.

Download from HuggingFace

# Install MLX CLI
pip install mlx-lm

# Download Model (Automatically Convert to MLX Format)
mlx_lm.convert --hf-path Qwen/Qwen2.5-7B-Instruct --mlx-path ~/models/qwen2.5-7b-mlx

# Quantization (Reducing Memory Usage)
mlx_lm.convert --hf-path Qwen/Qwen2.5-7B-Instruct --mlx-path ~/models/qwen2.5-7b-q4 -q

Recommended Models

ModelMLX SizeMemoryUse Case
Qwen 2.5 7B Q44.5GB8GBEveryday Q&A
Qwen 2.5 32B Q418GB24GBCode generation
Qwen 2.5 72B Q440GB48GBDeep analysis
DeepSeek R1 70B Q440GB48GBReasoning
Llama 3.1 70B Q440GB48GBGeneral purpose

Configuration and Startup

In the oMLX app:

  1. Settings → Models Directory: point to ~/models/
  2. Select Model: select a downloaded model from the dropdown menu
  3. Adjust Parameters:
    • Temperature: 0.7 (creative) → 0.1 (precise)
    • Top-P: 0.9
    • Max Tokens: 4096
    • Context Length: 32768

API Server Mode

oMLX provides an OpenAI-compatible API:

# Start API Server
mlx_lm.server --model ~/models/qwen2.5-7b-mlx --port 8080

Then use it in Claude Code:

{
  "apiKey": "not-needed",
  "baseURL": "http://localhost:8080/v1",
  "model": "qwen2.5-7b-mlx"
}

Performance Comparison (M3 Ultra)

ModeloMLX (MLX)Ollama (llama.cpp)Difference
Qwen 2.5 7B Q445 t/s38 t/s+18%
Qwen 2.5 32B Q418 t/s14 t/s+29%
Qwen 2.5 72B Q48 t/s6 t/s+33%

MLX has a more pronounced advantage on large models because it uses the GPU more effectively.


oMLX vs Ollama vs LM Studio

FeatureoMLXOllamaLM Studio
Underlying frameworkApple MLXllama.cppllama.cpp + MLX
API compatibility✅ OpenAI✅ OpenAI✅ OpenAI
GUI✅❌✅
Model formatMLX, safetensorsGGUFGGUF, MLX
Metal accelerationNativeVia MetalHybrid
Speed (large models)FastestModerateModerate

Frequently Asked Questions

Are MLX models taking up too much space?

Use quantization:

mlx_lm.convert --hf-path model-name --mlx-path output-dir -q --q-bits 4

Can't find a model in MLX format?

Search for mlx-community on HuggingFace; most popular models have MLX versions.

oMLX vs. using the MLX CLI directly?

oMLX provides a GUI and more convenient model management, making it suitable for everyday use. MLX CLI is better suited for scripting and automation.


Recommended Reading