Turning a few Macs into an AI cluster

One Mac can run a capable model and a whole agent loop locally. But models keep growing, contexts get longer and tasks get harder, and sooner or later a single machine runs out of memory, compute or bandwidth. MLX can spread the work across several Macs on your desk, for inference and for fine-tuning, with very few changes to the commands you already use.

The stack that makes it possible

  • Thunderbolt 5 cables are the physical link between machines.
  • RDMA over Thunderbolt (from macOS 26.2) moves data straight from one Mac's memory to another's, avoiding most CPU and operating-system overhead. That's the high bandwidth and low latency distributed work needs.
  • JACCL is an open-source collective communication library built on RDMA. It provides the primitives for sending data between Macs and combining results across the group, without you managing the transport. It isn't limited to machine learning; any distributed workload on Apple silicon can use it.
  • MLX and MLX-LM sit on top: they shard models, coordinate the machines and run inference and fine-tuning.

Mesh or ring

Communication time has two parts. Latency is a fixed cost for every operation, whatever the size. Transfer time grows with the amount of data and depends on bandwidth. Small messages are dominated by latency, large ones by bandwidth, and the wiring decides which you're good at.

A full mesh connects every Mac to every other, so any machine is one hop away: the lowest latency. A ring connects each Mac to its two neighbours, which needs fewer cables and ports and scales to more machines; the spare Thunderbolt ports can carry two or three cables per neighbour for extra bandwidth. If you wire a mesh, JACCL picks the best route for each operation automatically: mesh when latency matters, ring when bandwidth does.

Setting up the cluster

  1. Connect the Macs with Thunderbolt 5 cables, as a mesh or a ring.
  2. Enable RDMA on each one: open Settings, search for RDMA, turn on "Enable RDMA over Thunderbolt", and restart.
  3. Generate a hostfile with mlx.distributed_config, giving it the hostnames and an output path. With --auto-setup it checks every Mac is reachable over SSH, probes the Thunderbolt ports to map which machines are connected, configures each link for RDMA and writes the file. Without that flag it prints the commands so you can review and run them yourself. Use --backend jaccl for a mesh or jaccl-ring for a ring.

The hostfile is a JSON array with one entry per Mac:

[
  {
    "ssh": "mac-1",
    "ips": ["<local network IP>"],
    "rdma": ["<RDMA device for each Thunderbolt peer>"]
  }
]

ssh is how the launcher reaches the machine, ips is used for the initial handshake over your local network, and rdma lists the device for each Thunderbolt connection. You can also add environment variables that get set on every Mac at launch; MLX_METAL_FAST_SYNCH=1 is worth it, because computation runs on the GPU while communication runs on the CPU, and it speeds up the hand-off between them.

From then on, mlx.launch does the orchestration. Run it from any machine with SSH access, a laptop for example, with the hostfile and the program to run. It connects to each Mac and starts the program, and after that the Macs talk directly over Thunderbolt. MLX and your scripts need to be installed on every Mac.

Bigger and faster models

To chat with a model across the cluster, wrap the usual single-Mac command with mlx.launch, pointing it at the path of mlx_lm.chat on the remote machines:

mlx.launch --hostfile hosts.json -- \
  /path/to/mlx_lm.chat --model <model> --max-tokens 4096

MLX-LM shards the model and coordinates the machines. Two things change:

  • Speed. On four M3 Ultra Macs, a 27-billion-parameter model generated tokens almost three times faster than on one. The exact gain depends on the model's size and architecture.
  • Size. A one-trillion-parameter model needs about a terabyte for its weights even at 8-bit precision. That doesn't fit on one M3 Ultra, but it fits across four.

Pipeline or tensor parallelism

There are two ways to split a model across machines, and they trade speed for traffic differently.

  • Pipeline parallelism splits the model by depth. Each Mac holds a block of layers, and data moves through the machines in turn. It doesn't make generation faster, since each token still passes through every block one after another, but the machines only exchange data at the boundaries. Turn it on with --pipeline; not every model supports it.
  • Tensor parallelism splits the model by width. Each Mac holds part of every layer, and all of them work on the same token at once, which is where the speed-up comes from. The cost is communication at every layer for every token, so low latency matters, and a mesh pays off. It's the default in MLX-LM.

Faster fine-tuning

Fine-tuning works through the training data in batches: compute gradients, update the weights, repeat. Across several Macs, MLX-LM uses data parallelism: every Mac holds a full copy of the model and trains on a different batch, and the gradients are averaged so every copy takes the same update. With N Macs you can get through the data up to N times faster.

The command is almost the same as on one Mac. Launch mlx_lm.lora through mlx.launch, and multiply the batch size by the number of machines so each still processes the same amount per step. Fine-tuning a 9-billion-parameter model went from about 180 tokens per second on one M3 Ultra to about 600 on four, more than three times faster. The data never leaves your machines.

From the command line to your own code

Everything above uses the command line, but the same pieces are there in Python, Swift and C++. In Python with MLX-LM, you initialise the distributed group, choose the kind of parallelism and load the model with sharded_load; after that it behaves like a single-device model. For finer control, MLX has lower-level tools such as shard_linear for splitting a layer, and collective operations such as a distributed sum across all Macs. JACCL can also be built on its own with a C++ API, for distributed work that has nothing to do with machine learning. MLX-LM's built-in server can serve a model across the cluster too.

The result: faster inference, models with a trillion parameters, and quicker fine-tuning, on hardware you own and with your data staying private.