Ollama taps Apple’s MLX framework to make local AI models faster on Macs

TL;DR AI
2 min readKey summary
Ollama’s new release integrates Apple’s MLX framework to speed up local model inference on Apple Silicon.
The update also adds support for NVIDIA’s NVFP4 format and improves caching and quantization to reduce latency.
MLX support in this release is currently limited to the Qwen3.5-35B-A3B model, while Ollama continues to run open-weight models locally.



