Local agents that don’t keep you waiting

Running a model on a Mac is easy now. Running a coding agent on one is harder, and the reason is less about raw speed than about waiting: the long pause before the first word every time the agent sends a big prompt. oMLX is an open-source model server built on MLX that tackles exactly that. It follows on from running AI agents entirely on your Mac and turning a few Macs into an AI cluster.

Why agents keep you waiting

Before a model writes anything, it has to process the whole prompt and build a cache of it, the KV cache. Chat is gentle on that cache: each turn adds a little to the end. Coding agents aren't. They rewrite and reshuffle their context constantly, and each time the cache built in memory no longer matches and gets thrown away. On a long context, rebuilding it can take tens of seconds, again and again through a session.

oMLX keeps the cache instead of throwing it away. Cache blocks are written to the SSD; the most-used ones stay in RAM and older ones move to disk, least recently used first. When the agent comes back to a prompt it has seen before, the matching blocks are restored, even after the server restarts, rather than recomputed. oMLX's own figures put time to first token on long contexts at under five seconds from the second turn, against 30 to 90 seconds with a memory-only cache.

What else it adds

oMLX sits on top of MLX-LM, so it runs the same models, and adds what you need to serve them day to day:

  • A native Mac app. A menu bar app to start, stop and monitor the server, plus a web dashboard for models, chat and live metrics.
  • Two kinds of API. An OpenAI-compatible endpoint and an Anthropic-compatible one, so it works with Claude Code, OpenCode, Pi, Cursor and anything else that speaks either. The dashboard generates the exact config command for each tool.
  • Continuous batching, so several agents or subagents can share the model without queuing.
  • Tool calling in the formats the main model families use, and several models loaded at once, including vision, embedding and reranking models.
  • Your existing models. It reads the standard Hugging Face cache and your LM Studio folder, so there's nothing to download again.

LM Studio remains a friendly way to start, and Ollama is adding MLX support, but for agents the persistent cache is the difference that matters most. A lighter app also leaves more memory for the model itself.

Setting it up

  1. Install it. Download the app from omlx.ai and drag it to Applications; a welcome screen walks you through choosing a model folder and starting the server. It needs Apple silicon and macOS 15 or later.
  2. Get a model. Find an MLX version on Hugging Face (the mlx-community organisation has quantised builds of most popular models) and paste its address into oMLX's downloader. For agentic coding, mixture-of-experts models are a good fit: large overall, but only part of the model works on each token, so they stay fast.
  3. Connect your agent. Open the dashboard, pick the model and copy the command it gives you for your tool. In tools that take a custom provider, add one with oMLX's base URL and the API key you set in oMLX.

If you prefer the command line, you can run it from source instead:

git clone https://github.com/jundot/omlx
cd omlx && pip install -e .
omlx serve --model-dir ~/models

It then listens on localhost:8000 for any OpenAI-compatible client. With a private network such as Tailscale, you can also reach it from your other devices.

Plan for context, not just the model

A model's file size is only part of the memory it needs. The context and its cache grow as you work: an 8-bit mixture-of-experts model of about 36 GB can use around 80 GB in total once the context is long. Many models can handle very long contexts (Qwen 3.6 accepts up to 262,144 tokens), but the setting is per model and needs a reload to change, so pick a model size that leaves room for the context you actually use.

The agent matters too. Some agents use context more freely than others, so on limited hardware a leaner one, such as OpenCode, stretches the same memory further.

Where it fits

MLX-LM is still the place for training and lower-level experiments. For day-to-day local agents on a Mac, a server that remembers what it has already read makes the difference between an assistant you wait for and one you work with.