Agentic Research

M3 Ultra 512GB: The Dream Workstation for Running Multi-Agent Frameworks

2026/05/0911 min readBryan Chan閱讀中文原文
TopicsApple SiliconLocal LLM

Hardware Configuration

  • Chip: Apple M3 Ultra
  • CPU: 32 cores (24 performance + 8 efficiency)
  • GPU: 80 cores
  • Neural Engine: 32 cores
  • Unified memory: 512GB
  • Memory bandwidth: 819 GB/s
  • SSD: 8TB

What 512GB of Unified Memory Actually Means

In a traditional architecture, the CPU and GPU each have their own separate memory, and data must be copied between them. Apple's unified memory architecture means the CPU and GPU share the same 512GB.

For AI workloads, this means:

Running Multiple LLMs at the Same Time

Currently resident models (all simultaneously in memory):
├── DeepSeek R1 671B Q4     ~380GB  (via llama.cpp)
├── Qwen 2.5 72B Q4          ~40GB
├── Llama 3.1 70B Q4         ~40GB
├── Mistral 8x22B Q4         ~45GB
└── CodeLlama 34B Q4         ~20GB
                            ─────
Remaining available:                    ~0GB  (just filled to capacity)

In practice, you would not run this many large models at once. Usually 2-3 resident models are enough, with the rest of the memory left for Agent frameworks and the operating system.

Memory Usage of Agent Frameworks

Hermes Agent      ~500MB  (Node.js runtime)
OpenClaw          ~800MB  (Node.js + Skills)
Claude Code        ~1GB   (per session × 3 max)
Operating System + Tools    ~8GB
─────────────────────────
Total              ~12GB

Image/Video Generation

ComfyUI + SDXL         ~16GB (MPS)
Wan2.2 Video Generation         ~24GB (MPS)
─────────────────────────────
Total                   ~40GB

Conclusion: one M3 Ultra 512GB can run 2-3 Agent frameworks + 1-2 local LLMs + image generation at the same time, all without any swap.


Comparison with Other Options

OptionMemoryCostNoiseSuited For
M3 Ultra 512GB512GB unified~$8,000FanlessThe ultimate personal setup
M2 Ultra 192GB192GB unified~$5,000FanlessAn advanced personal setup
RTX 4090 × 496GB VRAM × 4~$8,000LoudTraining/clusters
MacBook Pro M3 Max128GB~$4,000NoneA mobile setup
Cloud A10080GB~$1/hrN/AAn elastic setup

Real-World Work Scenarios

Scenario 1: Everyday Development

Running:
Ollama (Qwen 2.5 32B)    ~18GB
  OpenClaw + Hermes          ~2GB
VS Code + Terminal             ~4GB
  Claude Code (1 session)    ~1GB
─────────────────────────────────
Total                       ~25GB / 512GB (only 5% used)

Completely imperceptible, with plenty of headroom left.

Scenario 2: Full-Load Stress Test

Running:
Ollama (DeepSeek R1 70B)  ~40GB
  Ollama (Qwen 72B)         ~40GB
ComfyUI + SDXL            ~16GB
Wan2.2 Video Generation    ~24GB
  OpenClaw + Hermes          ~2GB
  Claude Code (3 sessions)   ~3GB
macOS + Tools              ~8GB
─────────────────────────────────
Total                        ~133GB / 512GB: only 26% used

Memory is not the bottleneck at all; GPU utilization is.


Shortcomings

  1. Limited GPU core count: an 80-core GPU is not as good as multiple RTX 4090s for large-batch training
  2. The MPS ecosystem is not mature enough: some PyTorch operations do not support MPS and need a CPU fallback
  3. Price: the 512GB version is not cheap, a considerable investment for an individual developer
  4. CUDA is unavailable: some frameworks (such as bitsandbytes) can only run on CUDA

Conclusion

The M3 Ultra 512GB is an excellent choice for running multi-Agent frameworks:

  • Unified memory means LLM inference is not limited by VRAM
  • Zero noise makes it suited to long-running operation
  • Extremely low power consumption (compared with x86 + GPU setups)
  • Enough to run an entire three-layer Agent framework + multiple local LLMs at the same time

If you take local AI development seriously, this machine can be called the ultimate option available today.