Exo: Turn multiple machines into an LLM cluster maison
Exo links multiple computers into a cluster to load a model larger than a single machine's memory, with automatic device discovery on the network. The real benefit is cumulative memory: multiple Mac Studios with 512 GB can load a model that none of them could support alone. The real cost is the network: without RDMA over Thunderbolt 5 (recent Macs running macOS 26.2), latency hurts speed, and Linux currently runs on CPU only.
Exo is an open-source project maintained by exo labs that connects multiple machines into an inference cluster, allowing it to load models larger than a single device can support. This guide distinguishes what is confirmed by the official repository from what is still only an announcement: the cumulative memory actually available, measured speed gains on Apple hardware, network cost between machines, and the precise matrix of supported platforms before investing in multiple devices.
#What Exo actually does
Exo (the exo-explore/exo repository, with nearly 48,000 GitHub stars as of September 28, 2026) is described by its creators as a tool to “run cutting-edge AI locally.” Devices running Exo discover each other automatically on the network, without manual configuration, and expose a dashboard and an API at http://localhost:52415 on each node.
The project highlights “topology-aware” tensor parallelism: Exo evaluates network topology in real time (latency and bandwidth between every pair of machines) and each device's resources to decide how to distribute a model, rather than applying the same fixed split regardless of the available hardware.
#Cumulative memory: the real contribution
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The most tangible benefit of an Exo cluster is the sum of the available memory. One configuration illustrated by the project itself combines 4 Mac Studio M3 Ultra systems with 512 GB each to load DeepSeek v3.1 in 8-bit and Kimi-K2-Thinking in 4-bit simultaneously—models that none of these machines could load alone, even with 512 GB.
In terms of compute, Exo claims a gain from tensor parallelism: up to 1.8x faster on 2 devices and 3.2x on 4 devices. These are speedup factors for distributed computation, not tokens-per-second throughput: the actual number of tokens generated per second always depends on the model, quantization, and—as we’ll discuss below—the network between the machines.
#Network cost: what RDMA changes
Connecting machines over a network introduces latency that a single GPU doesn't have. Exo addresses this with RDMA support (remote direct memory access) over Thunderbolt 5, announced when the feature was released, claiming a 99% reduction in latency between devices compared with a standard network connection.
This RDMA capability is not universal: it relies on a feature added in macOS 26.2 and works only on Macs equipped with Thunderbolt 5 — Mac mini M4 Pro, Mac Studio M4 Max or M3 Ultra, MacBook Pro M4 Max according to the official list. The project also requires all devices in the RDMA cluster to be interconnected, with TB5-certified cables, and the macOS version (including beta versions) to be strictly identical on every machine; otherwise, the RDMA ports may fail to discover one another.
#Matrix of platforms that are actually supported
| Platform | Acceleration | Status |
|---|---|---|
| macOS (Apple Silicon) | GPU via MLX | Primary path, Thunderbolt 5 RDMA available on recent Macs running macOS 26.2+ |
| Linux (x86/ARM) | CPU only | GPU support is under development according to the official README |
| Windows | No official mention | Not documented in the repository as of September 28, 2026 |
This table contradicts a common assumption: Exo is not a tool that runs a heterogeneous Mac plus PC cluster at full speed. The MLX inference backend is specific to Apple Silicon; on Linux, the official README clearly states that Exo currently runs on the CPU, without GPU acceleration, with GPU support announced as “in development” and no delivery date specified.
#Deploying a cluster in practice
- 01Prepare each macOS machineInstall Xcode (for the Metal toolchain), Homebrew, uv, and Node—the documented prerequisites for building and running Exo from source on macOS.
- 02Clone and run Exogit clone du dépôt, construction du tableau de bord (npm install && npm run build), puis uv sync --extra mlx et uv run exo sur chaque machine.
- 03Check automatic discoveryDevices on the same network discover one another without configuration; the dashboard available at http://localhost:52415 should list every active cluster node.
- 04Enable RDMA if the hardware supports itOn Thunderbolt 5 Macs running macOS 26.2+, enable rdma_ctl enable from Recovery mode on each machine, then connect all the devices to one another with TB5-certified cables.
- 05Load a sharded modelFrom the dashboard or API, choose a model whose size exceeds the memory of a single device to verify that partitioning works in practice before targeting even larger models.
#Compatible APIs and existing clients
Exo exposes several compatible APIs: OpenAI Chat Completions, OpenAI Responses, Claude Messages, and Ollama—a client already written for one of these formats can point to the Exo cluster without rewriting. It’s a pragmatic choice that avoids having to adopt a proprietary protocol to use the cluster.
Environment variables supplement the configuration for advanced use: EXO_OFFLINE to run without an internet connection once the models are already cached, EXO_MODELS_READ_ONLY_DIRS to share read-only model storage between machines (useful for a shared NFS mount), and EXO_LIBP2P_NAMESPACE to isolate multiple Exo clusters on the same physical network.
#Custom models from Hugging Face
Exo is not limited to a closed catalog of models: the project supports custom models directly from the Hugging Face Hub, beyond the demonstration deployments highlighted by its creators (DeepSeek v3.1 with 671 billion parameters, Qwen3-235B, Kimi-K2-Thinking). This makes it possible to test a recent model as soon as it is published on the Hub, without waiting for it to be added to an officially supported list.
This openness has a documented tradeoff: custom models that require trust_remote_code in their configuration must be explicitly enabled, as this setting remains disabled by default for security reasons. Running arbitrary code supplied by a third-party Hugging Face repository across all machines in the cluster is therefore not the default behavior, but a choice to make consciously for models that require it.
#Security: a cluster with no documented authentication
The official README and dashboard documentation mention no authentication mechanism, API key, token, or TLS encryption for the Exo API and dashboard: as documented on September 28, 2026, anyone who can reach port 52415 on a node can query the API or view the dashboard without a login step.
This point deserves even more attention because each machine in the cluster exposes its own API on the local network: in a multi-device setup, the exposed surface is that of each node individually, not just that of a single entry point that could be protected in isolation.
#Troubleshooting: symptoms, cause, fix
| Symptom | Likely cause | Correction |
|---|---|---|
| RDMA ports are not discoverable between two Macs | Different macOS versions across machines, including beta versions | Strictly align the macOS version (including beta versions) across every device in the cluster |
| A custom model refuses to load from Hugging Face | trust_remote_code required by the model but disabled by default | Explicitly enable trust_remote_code for this specific model, understanding the code it executes |
| Disappointing throughput despite adding several machines | No RDMA: the cluster runs over Wi-Fi or standard Ethernet, making it more latency-sensitive | Verify RDMA eligibility (Thunderbolt 5, macOS 26.2+) or move the machines closer together on a dedicated wired network |
| A Linux machine in the cluster does not appear to accelerate inference | GPU support on Linux is still under development; the machine contributes CPU only | Count this machine only as memory and CPU compute, not as a GPU accelerator |
#Limits and pitfalls to anticipate
- No documented Windows support
- The official repository mentions no Windows support as of September 28, 2026; the installation paths cover macOS and Linux only.
- Linux limited to the CPU
- GPU-accelerated inference on Linux is in development, with no announced release date; a Linux-only cluster will be significantly slower than an equivalent Apple Silicon cluster.
- Hardware-intensive RDMA
- Thunderbolt 5, macOS 26.2 or later, and strictly identical system versions on every machine: a single device that fails any of these requirements falls back to a slower conventional network.
- Throughput not guaranteed
- The acceleration factors (1.8x, 3.2x) measure the benefit of tensor parallelism for distributed computation, not a tokens-per-second throughput that can be directly transferred to your own model and network.
- Multi-GPU LLM with llama.cpp: tensor-split 2× RTX 3090
- llama.cpp vs vLLM vs Exllama
- Deploy an LLM in production with Docker Compose
- Security principles for an exposed inference server
- Source: the official Exo GitHub repository
- Source: official README of the Exo repository
Can Exo Mix Macs and Windows PCs in the Same Cluster?+
Does Exo use the GPU on Linux?+
Do you need RDMA over Thunderbolt to use Exo?+
How much memory can an Exo cluster pool?+
Are the speed gains advertised by Exo (1.8x, 3.2x) guaranteed on my hardware?+
Are Exo’s dashboard and API protected by a password?+
Can Exo load any model published on Hugging Face?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.