Agentic Research

LM Studio Installation and Usage Guide: GUI Local LLM Management

2026/05/1012 min readBryan Chan閱讀中文原文
TopicsLocal LLMMLXMac

What is LM Studio?

LM Studio is a desktop GUI application that lets you download, manage, and run LLMs locally, without a command line. It supports models in GGUF and MLX formats.

  • Graphical interface: No terminal operations required; run models with a click
  • Built-in model download: Search and download directly from HuggingFace
  • OpenAI-compatible API: Automatically provides a localhost:1234/v1 endpoint after startup
  • Supports Apple Silicon GPU: Native acceleration for M-series chips

Installation

# Official Website Download
# https://lmstudio.ai/

# Or Homebrew
brew install --cask lm-studio

Model Download and Management

Steps

  1. Open LM Studio → click the Search icon on the left
  2. Search for a model, for example qwen2.5-coder-7b
  3. Select the GGUF format (Q4_K_M quantization is the most balanced)
  4. Click Download

Recommended Models

ModelSizeRecommended QuantizationUse Case
Qwen 2.5 Coder 7B4.7GBQ4_K_MCode
Qwen 2.5 32B19GBQ4_K_MGeneral
DeepSeek Coder V215GBQ4_K_MCode
Llama 3.1 8B4.9GBQ4_K_MEnglish
Mistral Nemo 12B7GBQ4_K_MFast

Running Models

  1. Click the Chat icon on the left.
  2. Select the downloaded model from the top dropdown menu.
  3. Adjust the parameters on the right:
    • GPU Offload: Set to Max (load everything onto the GPU)
    • Context Length: 8192 (daily use) / 32768 (long documents)
    • Temperature: 0.7
  4. Once the model is loaded, you can start chatting.

API Server Mode

Starting the Server

Click Local Server on the left, then Start Server

Default port: http://localhost:1234

Using in Claude Code

{
  "apiKey": "not-needed",
  "baseURL": "http://localhost:1234/v1",
  "model": "qwen2.5-coder-7b-instruct"
}

Calling from Python

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:1234/v1",
    api_key="not-needed"
)

response = client.chat.completions.create(
    model="qwen2.5-coder-7b-instruct",
    messages=[{"role": "user", "content": "Write a Python function to sort a list"}]
)

print(response.choices[0].message.content)

MLX vs GGUF Format Selection

LM Studio supports both GGUF and MLX formats:

FeatureGGUFMLX
Speed (M3 Ultra)FastFaster
Model library breadthVery extensiveLess extensive
Quantization optionsQ2-Q8, K-quantsQ4, Q8
Cross-platform✅Apple Silicon only
Recommended use casesGeneral purposeMaximizing performance on M-series chips

Recommendation: When an MLX version is available, prefer MLX; otherwise, use GGUF Q4_K_M.


Advanced Configuration

Running Multiple Models Simultaneously

# Terminal 1: General Model
curl http://localhost:1234/v1/chat/completions -d '{"model":"qwen2.5-32b","messages":[...]}'
# Terminal 2: Code Model (Different Port)
# Start a second server in LM Studio, with the port changed to 1235
curl http://localhost:1235/v1/chat/completions -d '{"model":"qwen2.5-coder-7b","messages":[...]}'

GPU Layer Adjustment

Adjust GPU Offload in the right panel of LM Studio:

  • Max: Load everything into GPU (fastest, requires sufficient VRAM/unified memory)
  • Auto: Automatic balancing
  • Manual: Specify the number of layers (suitable for running multiple models simultaneously)

LM Studio vs oMLX vs Ollama Comparison

FeatureLM StudiooMLXOllama
InterfaceGUIGUICLI
Model downloadBuilt-inHuggingFaceollama pull
Model formatGGUF + MLXMLXGGUF
API compatibility✅✅✅
GPU acceleration✅✅ (MLX)✅ (Metal)
Beginner friendliness⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐⭐

Recommended Reading