oMLX Installation and Usage Guide: Running MLX Models Natively on Apple Silicon
TopicsoMLXMLXApple SiliconLocal LLMMac
What is oMLX?
oMLX is a desktop application based on the Apple MLX framework that lets you natively run MLX-format LLMs on Apple Silicon Macs. MLX is Apple's official machine learning framework, optimized specifically for M-series chips.
- Native Apple Silicon: Leverages the M3 Ultra's GPU + Neural Engine
- Unified Memory: Models load directly into 512GB unified memory
- No Rosetta required: Native ARM64, faster than llama.cpp
Installation
# Download the Latest Version
# https://github.com/open-mlx/oMLX/releases
# Or via Homebrew (if there is a cask)
brew install --cask omlx
After downloading the .dmg, drag it into Applications.
Model Import
oMLX supports .safetensors and MLX format models.
Download from HuggingFace
# Install MLX CLI
pip install mlx-lm
# Download Model (Automatically Convert to MLX Format)
mlx_lm.convert --hf-path Qwen/Qwen2.5-7B-Instruct --mlx-path ~/models/qwen2.5-7b-mlx
# Quantization (Reducing Memory Usage)
mlx_lm.convert --hf-path Qwen/Qwen2.5-7B-Instruct --mlx-path ~/models/qwen2.5-7b-q4 -q
Recommended Models
| Model | MLX Size | Memory | Use Case |
|---|---|---|---|
| Qwen 2.5 7B Q4 | 4.5GB | 8GB | Everyday Q&A |
| Qwen 2.5 32B Q4 | 18GB | 24GB | Code generation |
| Qwen 2.5 72B Q4 | 40GB | 48GB | Deep analysis |
| DeepSeek R1 70B Q4 | 40GB | 48GB | Reasoning |
| Llama 3.1 70B Q4 | 40GB | 48GB | General purpose |
Configuration and Startup
In the oMLX app:
- Settings → Models Directory: point to
~/models/ - Select Model: select a downloaded model from the dropdown menu
- Adjust Parameters:
- Temperature: 0.7 (creative) → 0.1 (precise)
- Top-P: 0.9
- Max Tokens: 4096
- Context Length: 32768
API Server Mode
oMLX provides an OpenAI-compatible API:
# Start API Server
mlx_lm.server --model ~/models/qwen2.5-7b-mlx --port 8080
Then use it in Claude Code:
{
"apiKey": "not-needed",
"baseURL": "http://localhost:8080/v1",
"model": "qwen2.5-7b-mlx"
}
Performance Comparison (M3 Ultra)
| Model | oMLX (MLX) | Ollama (llama.cpp) | Difference |
|---|---|---|---|
| Qwen 2.5 7B Q4 | 45 t/s | 38 t/s | +18% |
| Qwen 2.5 32B Q4 | 18 t/s | 14 t/s | +29% |
| Qwen 2.5 72B Q4 | 8 t/s | 6 t/s | +33% |
MLX has a more pronounced advantage on large models because it uses the GPU more effectively.
oMLX vs Ollama vs LM Studio
| Feature | oMLX | Ollama | LM Studio |
|---|---|---|---|
| Underlying framework | Apple MLX | llama.cpp | llama.cpp + MLX |
| API compatibility | ✅ OpenAI | ✅ OpenAI | ✅ OpenAI |
| GUI | ✅ | ❌ | ✅ |
| Model format | MLX, safetensors | GGUF | GGUF, MLX |
| Metal acceleration | Native | Via Metal | Hybrid |
| Speed (large models) | Fastest | Moderate | Moderate |
Frequently Asked Questions
Are MLX models taking up too much space?
Use quantization:
mlx_lm.convert --hf-path model-name --mlx-path output-dir -q --q-bits 4
Can't find a model in MLX format?
Search for mlx-community on HuggingFace; most popular models have MLX versions.
oMLX vs. using the MLX CLI directly?
oMLX provides a GUI and more convenient model management, making it suitable for everyday use. MLX CLI is better suited for scripting and automation.
Recommended Reading
More in Tools
- PaddleOCR in Practice: Extracting Hong Kong Stock Annual Report Financial Data in 83 Seconds
- Webb-Site: The Essential Hidden Treasure for Hong Kong Stock Research, a One-Click Tool to Get Annual Report PDFs for All Listed Companies
- Academic Research Skills Deep Technical Breakdown: How 45+ Agents Collaborate to Complete the Full Workflow from Literature Review to Peer Review
- AI Engineering from Scratch Deep Dive: 435 Lessons × 20 Stages