Ollama: The Best Way to Run LLMs Locally
Why Choose Ollama?
- Fully local operation: Data never leaves your machine, ensuring privacy and security
- One-command installation: Just run
brew install ollama - Rich model selection: Llama 3, Mistral, Qwen 2.5, DeepSeek, Gemma, and more
- API compatibility: Offers an OpenAI-compatible API for seamless migration
- Mac optimization: Supports Metal GPU acceleration
Installation
# macOS
brew install ollama
# Linux
curl -fsSL https://ollama.com/install.sh | sh
Start the service:
ollama serve
Common Model Recommendations
| Model | Size | Best For |
|---|---|---|
deepseek-r1:8b | 4.9GB | Reasoning + code |
qwen2.5:7b | 4.7GB | General Chinese and English |
llama3.1:8b | 4.9GB | General English |
mistral:7b | 4.1GB | Lightweight and fast |
codellama:7b | 3.8GB | Code only |
Basic Usage
# Pull Model
ollama pull deepseek-r1:8b
# Interactive chat
ollama run deepseek-r1:8b
# Single-turn Q&A
ollama run deepseek-r1:8b "Explain quantum computing in simple terms"
API Mode
Ollama provides a REST API at http://localhost:11434 by default:
curl http://localhost:11434/api/generate -d '{
"model": "deepseek-r1:8b",
"prompt": "Why is the sky blue?",
"stream": false
}'
Using Ollama in Claude Code
Configure ~/.claude.json or environment variables to point to local Ollama:
{
"apiKey": "ollama",
"baseURL": "http://localhost:11434/v1"
}
Real-World Use Cases
Scenario 1: Local Code Assistant
On an airplane with no network, use Ollama to run qwen2.5-coder:7b as a local code assistant. Configure the Continue extension for VS Code to point to http://localhost:11434/v1, and you can use code completion and explanation features offline, without uploading your private code to the cloud at all.
Scenario 2: Batch Documentation Generation
Need to generate OpenAPI specifications for 50 API endpoints? Write a script to call Ollama's /v1/chat/completions endpoint in batch, passing the code for one endpoint in each request to automatically generate documentation. Running locally incurs no API fees, and no data leaves your environment.
Advanced Configuration: Custom Modelfile
# Create a custom model configuration
ollama pull qwen2.5:7b
ollama show qwen2.5:7b --modelfile > Modelfile
# Edit the Modelfile and add a system prompt
cat > Modelfile << 'EOF'
FROM qwen2.5:7b
SYSTEM "You are a professional programming assistant. When answering, you must: 1. Give the conclusion first 2. Provide code examples 3. Explain the key steps. Use Traditional Chinese."
PARAMETER temperature 0.3
PARAMETER num_ctx 4096
EOF
# Create the custom model
ollama create qwen-dev -f Modelfile
ollama run qwen-dev
Notes
- 7B models require at least 8GB RAM, and 13B models require 16GB
- Metal acceleration performs very well on Mac M-series chips
- Pulling a model for the first time requires downloading several GB, so be patient
- It is recommended to use
ollama listto manage downloaded models - Use
ollama rm <model>to delete unused models and free up space
Recommended Reading
- Ollama official documentation, complete usage guide
- OpenAI API compatibility documentation, API reference for comparison
- LM Studio, another local LLM runtime option