Running AI agents entirely on your Mac
A chat model answers a question. An agent gets work done: it asks the model what to do next, runs a command or reads a file, looks at the result, and goes round again until the task is finished. MLX can run that whole loop on a Mac, with no cloud and no API keys.
That has three practical upsides: your code and data stay on your machine, it works offline, and there's no per-token bill.
From chat to the agentic loop
With chat, you send a prompt, get an answer, and acting on it is your job. With an agent, the agent talks to the model to decide the next step, then calls tools to carry it out: running commands, reading files, calling APIs. It observes the results and asks the model again.
For example, a local agent can fetch recent pull requests from a GitHub repository, read the diffs and summarise what needs attention. The model runs on the Mac; only the git commands touch the network.
The local stack, in four layers
- MLX. Apple's open-source array framework for Apple silicon. It handles the low-level maths, Metal acceleration and memory.
- MLX-LM. Loads, runs, quantises and fine-tunes language models, with support for thousands of models from Hugging Face.
- MLX-LM Server. A persistent HTTP server that exposes the model through an OpenAI-compatible API, with structured tool calling and support for reasoning models. It's a drop-in replacement for a cloud API.
- The agent. Anything that speaks the OpenAI chat completions protocol: Xcode, OpenCode, Pi, or your own script.
Popular tools such as Ollama, LM Studio and vLLM also build on MLX, so if you already use one of them on a Mac, you may be running on it.
Set it up in three steps
- Install MLX-LM.
pip install mlx-lm- Start the server with a model that supports tool calling. Start small to check the setup works.
mlx_lm.server --model <model-with-tool-calling>- Point your agent at it. In most agent tools, set the provider's base URL to the local server (port 8080 by default) and choose the model name the server expects. In Xcode, add a locally hosted chat provider in the Intelligence settings and enter the same port.
The agent neither knows nor cares that the model is on your Mac rather than in a data centre.
What makes it fast enough
- Reading context. Agentic sessions often run to hundreds of thousands of tokens, and most of them are read, not written: every tool result has to be processed before the next step. The M5's Neural Accelerators make matrix multiplication four times faster than on the M4, and MLX's kernels turn that into almost four times faster prompt processing. It needs no code changes; MLX picks the best kernel for the hardware.
- Subagents at once. Agents often split work across subagents, one reading docs, one searching code, one writing tests. MLX-LM Server uses continuous batching: requests are grouped on the GPU, and new ones join a batch already in progress instead of waiting in a queue.
- Models too big for one Mac. Some models don't fit even in 512 GB; a recent 1.6-trillion-parameter model needs more than 800 GB for its weights alone. MLX can spread a model across several Macs over Thunderbolt or Ethernet, which also splits the prompt processing. With Thunderbolt RDMA from macOS 26.2, that's up to three times faster on four Macs. You start it with
mlx.launchand a hostfile listing the machines.
Real work, fully local
- A new app from scratch. Starting from an empty Xcode project, a local agent can plan and write a working SwiftUI app, building it with xcodebuild and fixing errors as it goes, then keep iterating on the changes you ask for.
- A bug fix inside Xcode. With Xcode connected to the same local server, the model can find a bug, read the surrounding code and write a fix you can build and run straight away.
Why it matters
Local agents change the trade-offs. Private code can stay private, a long agent session doesn't run up a bill, and you can work on a plane. The catch is capability: local models are smaller than the largest cloud ones, so pick tasks that fit, and treat multi-Mac setups as the route to bigger models.
Everything in the stack is open source and available now. To try it: install MLX-LM, launch the server, and point your favourite agent at it.