Comparison of Other Inference Frameworks: A Complete Guide to llama.cpp, MLX, and ExLlamaV2
Overview of Three Major Inference Frameworks
| Feature | llama.cpp | MLX | ExLlamaV2 |
|---|---|---|---|
| Platform | All platforms | Apple Silicon only | NVIDIA GPU only |
| Format | GGUF | MLX / safetensors | GPTQ / EXL2 |
| Speed (M3 Ultra) | Medium | Fastest | Not supported |
| Quantization | Q2-Q8, K-quants | Q4, Q8 | 2-8 bit |
| Ecosystem | Most mature (Ollama) | Apple official | High-end GPU users |
llama.cpp
Ollama's underlying engine, supporting CPU/GPU hybrid inference on all platforms.
git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make -j
# Running After Downloading a GGUF Model
./llama-cli -m model.gguf -p "Hello world" -ngl 99
- Advantages: Largest ecosystem, Ollama integration, richest collection of GGUF models
- Disadvantages: Not as fast as MLX on Apple Silicon
MLX (Apple)
Officially released by Apple, optimized specifically for M-series chips.
pip install mlx-lm
mlx_lm.generate --model mlx-community/Qwen2.5-7B-Instruct-4bit --prompt "Hello"
- Advantages: Fastest on Apple Silicon, most efficient use of unified memory
- Disadvantages: Apple Silicon only, smaller model library
ExLlamaV2
A high-performance inference engine designed specifically for NVIDIA GPUs.
git clone https://github.com/turboderp/exllamav2
pip install exllamav2
- Advantages: Fastest on NVIDIA GPUs, supports GPTQ
- Disadvantages: NVIDIA GPUs only, higher barrier to entry
Selection Recommendations
| Your environment | Recommendation |
|---|---|
| Mac M-series | MLX (fastest) > llama.cpp (Ollama) |
| NVIDIA GPU | ExLlamaV2 > llama.cpp |
| CPU only | llama.cpp |
| Prioritize simplicity | Ollama (built on llama.cpp) |
| Prioritize maximum performance | MLX (Mac) / ExLlamaV2 (NVIDIA) |
Model Formats and Quantization Trade-offs
The three frameworks are each tied to different formats, so choosing a framework is to some extent choosing a format. It is worth thinking this through before downloading models.
GGUF (llama.cpp): A single file carries its own metadata, supports mixed CPU and GPU offloading, and K-quants provide multiple compression levels. Its advantages are convenient cross-platform distribution and the widest range of model sources; its disadvantage is that file naming and quantization labels for the same model are relatively complex, so you need to confirm which version you downloaded.
MLX / safetensors (MLX): Aligned with Apple's unified memory architecture, its quantization is mainly low-bit, and its ecosystem revolves around community-converted MLX versions. If you only have safetensors on hand, you usually need to perform a conversion first.
GPTQ / EXL2 (ExLlamaV2): Targeted at NVIDIA GPUs, EXL2's flexibility lies in allocating different bit widths to different layers, trading off between VRAM and quality.
| Trade-off | Lower-bit quantization | Higher-bit quantization |
|---|---|---|
| Memory usage | Lower | Higher |
| Generation speed | Usually faster | Slower |
| Output quality | May decrease | Closer to original |
| Use case | Memory-constrained | Ample memory |
A practical recommendation is to start with medium quantization and run the entire workflow end to end, confirming that it loads and responds correctly, then adjust according to remaining memory and quality requirements. Do not pursue the lowest bit width from the start, otherwise you may mix quality issues with configuration issues and find it hard to determine the cause. Another common pitfall is format incompatibility: once you have GGUF, you can only use it in frameworks that support GGUF, and switching frameworks requires conversion, so be sure to keep the original weights before converting.
Key Tuning Points for Each Framework
llama.cpp: Core parameters include the number of layers offloaded to the GPU, thread count, batch size, and context length. Increasing the layer count makes it rely more on the GPU; if memory is insufficient, reduce it. Context length directly affects memory usage; setting it too high will cause failure even if the model itself fits. For long-term use, it is recommended to launch in server mode, which provides an OpenAI-compatible interface, making it easy for other tools to call.
MLX: The model must be in MLX format. If only raw weights are available, they can be converted first. Because unified memory lets the model and context share the same pool, long contexts require sufficient headroom, otherwise failures are likely to occur midway through generation. Using a version already quantized by the community can save the conversion step, but pay attention to whether the quantization method matches your needs.
ExLlamaV2: In most cases, it is used through the Python API or example scripts, and the model must first be converted to the corresponding format. VRAM size determines the bit configuration available per layer; the tighter the VRAM, the more you must compromise on bit width. This setup is suitable for users who want low latency on a single GPU and are willing to handle the conversion process.
The common principle across the three frameworks is: first validate the workflow with a small model, then tune parameters with the target model; record the result each time you change a parameter, avoiding changing multiple variables at once and being unable to attribute the effect.
Performance Validation Method
Performance comparisons must be made on the same basis; otherwise, the conclusions are meaningless.
- Hold the prompt content and output length fixed, and measure generation speed and first-token latency separately.
- Change only one variable at a time (quantization level, number of offloaded layers, context length) and compare the differences.
- Monitor peak memory and confirm that it does not spill into swap. Once it does, the speed data is unusable.
- Pay attention to cooling and throttling: in long continuous tests, later results will be lower than earlier ones, so use the same conditions when comparing.
- When comparing across frameworks, use equivalent quantized versions of the same model and run them with the same prompt.
It is recommended to script the tests and fix the inputs and measurement method, so that when changing versions or parameters, you can quickly rerun and compare.
Common Mistakes and Checklist
Common Mistakes
- Comparing speeds across different quantization versions, which invalidates the conclusion.
- Looking only at generation speed and ignoring first-token latency, when the interactive experience is actually poor.
- Expecting ExLlamaV2 to be available on Apple Silicon, or conversely expecting MLX on NVIDIA.
- Claiming usability even after the model has spilled into swap, when actual speed and stability are unacceptable.
- Ignoring the impact of context length on memory, only to find that loading fails after increasing it.
- Mixing quantization files from different frameworks, which causes immediate load errors.
Checklist
- Confirmed the framework support scope for the target platform.
- Confirmed that the model format matches the framework, and completed conversion when necessary.
- Selected the quantization level based on remaining memory.
- Recorded the combination of number of offloaded layers, number of threads, and context length.
- Measured generation speed and first-token latency.
- Confirmed that swap was not triggered during testing.
- Scripted the test workflow to facilitate rerunning and comparison.
Recommended Reading
More in Tools
- PaddleOCR in Practice: Extracting Hong Kong Stock Annual Report Financial Data in 83 Seconds
- Webb-Site: The Essential Hidden Treasure for Hong Kong Stock Research, a One-Click Tool to Get Annual Report PDFs for All Listed Companies
- Academic Research Skills Deep Technical Breakdown: How 45+ Agents Collaborate to Complete the Full Workflow from Literature Review to Peer Review
- AI Engineering from Scratch Deep Dive: 435 Lessons × 20 Stages