Agentic Research

oMLX + Qwen3 Coder Next Local Deployment Guide (Replacing Ollama)

2026/05/108 min readBryan Chan閱讀中文原文
TopicsoMLXLocal LLMApple SiliconOllama

Why Abandon Ollama?

Ollama was once the go-to for local LLMs, but recent versions have had issues:

  • Some models are forced to run in the cloud: models that should be local are routed to the Ollama cloud service, violating the original intent of "running locally"
  • Uncontrollable update strategy: automatic updates can change model behavior
  • Privacy uncertainty: it is unclear what data is sent to Ollama servers

Therefore, switch to oMLX, a desktop application built on Apple's native MLX framework: truly local, with zero cloud dependency.


Advantages of oMLX

FeatureOllamaoMLX
Frameworkllama.cpp (via Metal)Apple MLX (native)
Cloud dependencySome models forced to run in the cloud❌ Fully offline
Speed (M3 Ultra)Moderate18-33% faster
API compatibility✅ OpenAI✅ OpenAI
Model formatGGUF onlyMLX + safetensors
PrivacyUncertain✅ 100% local

Installation

# Download oMLX
# https://github.com/open-mlx/oMLX/releases
brew install --cask omlx

# Install MLX CLI
pip install mlx-lm

Deploy Qwen3 Coder Next

Download and Convert

# Method 1: Download directly from mlx-community (Recommended)
# Open oMLX → Search → search for "Qwen3-Coder-Next-MLX"

# Method 2: Manual Download and Conversion
mlx_lm.convert \
  --hf-path Qwen/Qwen3-Coder-Next \
  --mlx-path ~/models/qwen3-coder-next \
  -q --q-bits 4

Start API Server

mlx_lm.server \
  --model ~/models/qwen3-coder-next \
  --port 8080 \
  --host 0.0.0.0

Claude Code Local Instance Configuration

// ~/.claude-local.json
{
  "apiKey": "not-needed",
  "baseURL": "http://localhost:8080/v1",
  "model": "qwen3-coder-next"
}
claude --config ~/.claude-local.json --cwd ~/private-projects

OpenClaw Local Configuration

{
  "models": {
    "providers": {
      "local-coder": {
        "api_key": "not-needed",
        "base_url": "http://localhost:8080/v1",
        "model": "qwen3-coder-next"
      }
    },
    "routing": {
      "rules": [
        {"pattern": "private|secret|key|password", "provider": "local-coder"}
      ]
    }
  }
}

Performance Data (M3 Ultra)

ModeloMLXOllama (No Longer Used)
Qwen3 Coder Next Q455 t/s~42 t/s (est.)
Qwen 2.5 32B Q418 t/s14 t/s
Qwen 2.5 72B Q48 t/s6 t/s

Related Articles