Agentic Research

Ollama: The Best Way to Run LLMs Locally

2026/04/2811 min readBryan Chan閱讀中文原文
TopicsOllamaLLMLocal LLMMac

Why Choose Ollama?

  • Fully local operation: Data never leaves your machine, ensuring privacy and security
  • One-command installation: Just run brew install ollama
  • Rich model selection: Llama 3, Mistral, Qwen 2.5, DeepSeek, Gemma, and more
  • API compatibility: Offers an OpenAI-compatible API for seamless migration
  • Mac optimization: Supports Metal GPU acceleration

Installation

# macOS
brew install ollama

# Linux
curl -fsSL https://ollama.com/install.sh | sh

Start the service:

ollama serve

Common Model Recommendations

ModelSizeBest For
deepseek-r1:8b4.9GBReasoning + code
qwen2.5:7b4.7GBGeneral Chinese and English
llama3.1:8b4.9GBGeneral English
mistral:7b4.1GBLightweight and fast
codellama:7b3.8GBCode only

Basic Usage

# Pull Model
ollama pull deepseek-r1:8b

# Interactive chat
ollama run deepseek-r1:8b

# Single-turn Q&A
ollama run deepseek-r1:8b "Explain quantum computing in simple terms"

API Mode

Ollama provides a REST API at http://localhost:11434 by default:

curl http://localhost:11434/api/generate -d '{
  "model": "deepseek-r1:8b",
  "prompt": "Why is the sky blue?",
  "stream": false
}'

Using Ollama in Claude Code

Configure ~/.claude.json or environment variables to point to local Ollama:

{
  "apiKey": "ollama",
  "baseURL": "http://localhost:11434/v1"
}

Real-World Use Cases

Scenario 1: Local Code Assistant On an airplane with no network, use Ollama to run qwen2.5-coder:7b as a local code assistant. Configure the Continue extension for VS Code to point to http://localhost:11434/v1, and you can use code completion and explanation features offline, without uploading your private code to the cloud at all.

Scenario 2: Batch Documentation Generation Need to generate OpenAPI specifications for 50 API endpoints? Write a script to call Ollama's /v1/chat/completions endpoint in batch, passing the code for one endpoint in each request to automatically generate documentation. Running locally incurs no API fees, and no data leaves your environment.

Advanced Configuration: Custom Modelfile

# Create a custom model configuration
ollama pull qwen2.5:7b
ollama show qwen2.5:7b --modelfile > Modelfile

# Edit the Modelfile and add a system prompt
cat > Modelfile << 'EOF'
FROM qwen2.5:7b
SYSTEM "You are a professional programming assistant. When answering, you must: 1. Give the conclusion first 2. Provide code examples 3. Explain the key steps. Use Traditional Chinese."
PARAMETER temperature 0.3
PARAMETER num_ctx 4096
EOF

# Create the custom model
ollama create qwen-dev -f Modelfile
ollama run qwen-dev

Notes

  • 7B models require at least 8GB RAM, and 13B models require 16GB
  • Metal acceleration performs very well on Mac M-series chips
  • Pulling a model for the first time requires downloading several GB, so be patient
  • It is recommended to use ollama list to manage downloaded models
  • Use ollama rm <model> to delete unused models and free up space

Recommended Reading