Agentic Research

Comparison of Other Inference Frameworks: A Complete Guide to llama.cpp, MLX, and ExLlamaV2

2026/05/1025 min readBryan Chan閱讀中文原文
TopicsMLXInferenceLocal LLM

Overview of Three Major Inference Frameworks

Featurellama.cppMLXExLlamaV2
PlatformAll platformsApple Silicon onlyNVIDIA GPU only
FormatGGUFMLX / safetensorsGPTQ / EXL2
Speed (M3 Ultra)MediumFastestNot supported
QuantizationQ2-Q8, K-quantsQ4, Q82-8 bit
EcosystemMost mature (Ollama)Apple officialHigh-end GPU users

llama.cpp

Ollama's underlying engine, supporting CPU/GPU hybrid inference on all platforms.

git clone https://github.com/ggerganov/llama.cpp
cd llama.cpp
make -j

# Running After Downloading a GGUF Model
./llama-cli -m model.gguf -p "Hello world" -ngl 99
  • Advantages: Largest ecosystem, Ollama integration, richest collection of GGUF models
  • Disadvantages: Not as fast as MLX on Apple Silicon

MLX (Apple)

Officially released by Apple, optimized specifically for M-series chips.

pip install mlx-lm
mlx_lm.generate --model mlx-community/Qwen2.5-7B-Instruct-4bit --prompt "Hello"
  • Advantages: Fastest on Apple Silicon, most efficient use of unified memory
  • Disadvantages: Apple Silicon only, smaller model library

ExLlamaV2

A high-performance inference engine designed specifically for NVIDIA GPUs.

git clone https://github.com/turboderp/exllamav2
pip install exllamav2
  • Advantages: Fastest on NVIDIA GPUs, supports GPTQ
  • Disadvantages: NVIDIA GPUs only, higher barrier to entry

Selection Recommendations

Your environmentRecommendation
Mac M-seriesMLX (fastest) > llama.cpp (Ollama)
NVIDIA GPUExLlamaV2 > llama.cpp
CPU onlyllama.cpp
Prioritize simplicityOllama (built on llama.cpp)
Prioritize maximum performanceMLX (Mac) / ExLlamaV2 (NVIDIA)

Model Formats and Quantization Trade-offs

The three frameworks are each tied to different formats, so choosing a framework is to some extent choosing a format. It is worth thinking this through before downloading models.

GGUF (llama.cpp): A single file carries its own metadata, supports mixed CPU and GPU offloading, and K-quants provide multiple compression levels. Its advantages are convenient cross-platform distribution and the widest range of model sources; its disadvantage is that file naming and quantization labels for the same model are relatively complex, so you need to confirm which version you downloaded.

MLX / safetensors (MLX): Aligned with Apple's unified memory architecture, its quantization is mainly low-bit, and its ecosystem revolves around community-converted MLX versions. If you only have safetensors on hand, you usually need to perform a conversion first.

GPTQ / EXL2 (ExLlamaV2): Targeted at NVIDIA GPUs, EXL2's flexibility lies in allocating different bit widths to different layers, trading off between VRAM and quality.

Trade-offLower-bit quantizationHigher-bit quantization
Memory usageLowerHigher
Generation speedUsually fasterSlower
Output qualityMay decreaseCloser to original
Use caseMemory-constrainedAmple memory

A practical recommendation is to start with medium quantization and run the entire workflow end to end, confirming that it loads and responds correctly, then adjust according to remaining memory and quality requirements. Do not pursue the lowest bit width from the start, otherwise you may mix quality issues with configuration issues and find it hard to determine the cause. Another common pitfall is format incompatibility: once you have GGUF, you can only use it in frameworks that support GGUF, and switching frameworks requires conversion, so be sure to keep the original weights before converting.


Key Tuning Points for Each Framework

llama.cpp: Core parameters include the number of layers offloaded to the GPU, thread count, batch size, and context length. Increasing the layer count makes it rely more on the GPU; if memory is insufficient, reduce it. Context length directly affects memory usage; setting it too high will cause failure even if the model itself fits. For long-term use, it is recommended to launch in server mode, which provides an OpenAI-compatible interface, making it easy for other tools to call.

MLX: The model must be in MLX format. If only raw weights are available, they can be converted first. Because unified memory lets the model and context share the same pool, long contexts require sufficient headroom, otherwise failures are likely to occur midway through generation. Using a version already quantized by the community can save the conversion step, but pay attention to whether the quantization method matches your needs.

ExLlamaV2: In most cases, it is used through the Python API or example scripts, and the model must first be converted to the corresponding format. VRAM size determines the bit configuration available per layer; the tighter the VRAM, the more you must compromise on bit width. This setup is suitable for users who want low latency on a single GPU and are willing to handle the conversion process.

The common principle across the three frameworks is: first validate the workflow with a small model, then tune parameters with the target model; record the result each time you change a parameter, avoiding changing multiple variables at once and being unable to attribute the effect.


Performance Validation Method

Performance comparisons must be made on the same basis; otherwise, the conclusions are meaningless.

  • Hold the prompt content and output length fixed, and measure generation speed and first-token latency separately.
  • Change only one variable at a time (quantization level, number of offloaded layers, context length) and compare the differences.
  • Monitor peak memory and confirm that it does not spill into swap. Once it does, the speed data is unusable.
  • Pay attention to cooling and throttling: in long continuous tests, later results will be lower than earlier ones, so use the same conditions when comparing.
  • When comparing across frameworks, use equivalent quantized versions of the same model and run them with the same prompt.

It is recommended to script the tests and fix the inputs and measurement method, so that when changing versions or parameters, you can quickly rerun and compare.


Common Mistakes and Checklist

Common Mistakes

  • Comparing speeds across different quantization versions, which invalidates the conclusion.
  • Looking only at generation speed and ignoring first-token latency, when the interactive experience is actually poor.
  • Expecting ExLlamaV2 to be available on Apple Silicon, or conversely expecting MLX on NVIDIA.
  • Claiming usability even after the model has spilled into swap, when actual speed and stability are unacceptable.
  • Ignoring the impact of context length on memory, only to find that loading fails after increasing it.
  • Mixing quantization files from different frameworks, which causes immediate load errors.

Checklist

  • Confirmed the framework support scope for the target platform.
  • Confirmed that the model format matches the framework, and completed conversion when necessary.
  • Selected the quantization level based on remaining memory.
  • Recorded the combination of number of offloaded layers, number of threads, and context length.
  • Measured generation speed and first-token latency.
  • Confirmed that swap was not triggered during testing.
  • Scripted the test workflow to facilitate rerunning and comparison.

Recommended Reading