Ollama
Ollama runs supported models on MLX by default on Apple Silicon
Ollama 0.40.0 turns on Apple's MLX framework by default for supported models on Apple Silicon Macs.

Ollama now runs supported models on MLX, Apple’s machine-learning framework for its own chips, by default on Apple Silicon Macs. Ollama says MLX makes models respond faster and use less memory.
Ollama released version 0.40.0 on Oct. 6, 2026, a general release that followed release candidates rc1 to rc6 published Oct. 4 to 6. Its release notes on GitHub are headed “Models run on MLX on Apple Silicon by default.” They say: “In this release, on Apple Silicon devices, model architectures supported by the MLX runtime will automatically run on MLX.”
Gemma, Qwen and the decision models switch over
The release notes name gemma4, qwen3.6 and qwen3.5 as supported. They also list the decision models Nimble, tev1, clef and clef-flash, and the embedding model embeddinggemma-2, as running on MLX. The notes end with “We will continue testing and enabling additional models.”
Ollama said in a Sept. 29 blog post on decision models that it planned MLX acceleration on Apple Silicon. Version 0.40.0 delivers it by default for the models the notes list.
MLX went from a preview for M5 chips to the default
Ollama first put MLX in preview on March 30, 2026. That blog post required at least 32GB of unified memory and was optimized for M5, M5 Pro and M5 Max chips. Its listed model was Qwen3.5-35B-A3B in NVFP4 format.
Ollama reported prefill of 1,810 tokens per second, up from 1,154 before, and decode of 112 tokens per second, up from 58. These are Ollama’s own figures, with no named third-party eval.
Ollama says MLX responds faster and uses less memory
Ollama’s June 11, 2026 blog post, “Ollama’s highest performance on Apple Silicon yet with MLX,” said the runtime is “leaning more heavily on Apple’s unified memory and the Metal-backed MLX framework.” Models, it said, “output higher quality responses, respond faster, and use less memory.”
The same post said the NVFP4 format generated about 20% faster than q4_K_M on a MacBook Pro with an M5 Max running Gemma 4 12B. It also said NVFP4 “roughly halves the quality loss of 4-bit quantization, relative to unquantized BF16,” measured by perplexity. Those are Ollama’s own figures too.
Analysis
We think the main change is that version 0.40.0 switches supported models to MLX automatically. A Mac user running one of the listed models does not have to choose it. The release notes list only some model families and say testing continues. We read that to mean models off the list still run the old way for now.
