oMLX + Qwen3 Coder Next Local Deployment Guide (Replacing Ollama)
TopicsoMLXLocal LLMApple SiliconOllama
Why Abandon Ollama?
Ollama was once the go-to for local LLMs, but recent versions have had issues:
- Some models are forced to run in the cloud: models that should be local are routed to the Ollama cloud service, violating the original intent of "running locally"
- Uncontrollable update strategy: automatic updates can change model behavior
- Privacy uncertainty: it is unclear what data is sent to Ollama servers
Therefore, switch to oMLX, a desktop application built on Apple's native MLX framework: truly local, with zero cloud dependency.
Advantages of oMLX
| Feature | Ollama | oMLX |
|---|---|---|
| Framework | llama.cpp (via Metal) | Apple MLX (native) |
| Cloud dependency | Some models forced to run in the cloud | ❌ Fully offline |
| Speed (M3 Ultra) | Moderate | 18-33% faster |
| API compatibility | ✅ OpenAI | ✅ OpenAI |
| Model format | GGUF only | MLX + safetensors |
| Privacy | Uncertain | ✅ 100% local |
Installation
# Download oMLX
# https://github.com/open-mlx/oMLX/releases
brew install --cask omlx
# Install MLX CLI
pip install mlx-lm
Deploy Qwen3 Coder Next
Download and Convert
# Method 1: Download directly from mlx-community (Recommended)
# Open oMLX → Search → search for "Qwen3-Coder-Next-MLX"
# Method 2: Manual Download and Conversion
mlx_lm.convert \
--hf-path Qwen/Qwen3-Coder-Next \
--mlx-path ~/models/qwen3-coder-next \
-q --q-bits 4
Start API Server
mlx_lm.server \
--model ~/models/qwen3-coder-next \
--port 8080 \
--host 0.0.0.0
Claude Code Local Instance Configuration
// ~/.claude-local.json
{
"apiKey": "not-needed",
"baseURL": "http://localhost:8080/v1",
"model": "qwen3-coder-next"
}
claude --config ~/.claude-local.json --cwd ~/private-projects
OpenClaw Local Configuration
{
"models": {
"providers": {
"local-coder": {
"api_key": "not-needed",
"base_url": "http://localhost:8080/v1",
"model": "qwen3-coder-next"
}
},
"routing": {
"rules": [
{"pattern": "private|secret|key|password", "provider": "local-coder"}
]
}
}
}
Performance Data (M3 Ultra)
| Model | oMLX | Ollama (No Longer Used) |
|---|---|---|
| Qwen3 Coder Next Q4 | 55 t/s | ~42 t/s (est.) |
| Qwen 2.5 32B Q4 | 18 t/s | 14 t/s |
| Qwen 2.5 72B Q4 | 8 t/s | 6 t/s |
Related Articles
More in Tools
- PaddleOCR in Practice: Extracting Hong Kong Stock Annual Report Financial Data in 83 Seconds
- Webb-Site: The Essential Hidden Treasure for Hong Kong Stock Research, a One-Click Tool to Get Annual Report PDFs for All Listed Companies
- Academic Research Skills Deep Technical Breakdown: How 45+ Agents Collaborate to Complete the Full Workflow from Literature Review to Peer Review
- AI Engineering from Scratch Deep Dive: 435 Lessons × 20 Stages