Advanced 14 minCluster

Exo: Turn multiple machines into an LLM cluster maison

Direct response

Exo links multiple computers into a cluster to load a model larger than a single machine's memory, with automatic device discovery on the network. The real benefit is cumulative memory: multiple Mac Studios with 512 GB can load a model that none of them could support alone. The real cost is the network: without RDMA over Thunderbolt 5 (recent Macs running macOS 26.2), latency hurts speed, and Linux currently runs on CPU only.

Exo is an open-source project maintained by exo labs that connects multiple machines into an inference cluster, allowing it to load models larger than a single device can support. This guide distinguishes what is confirmed by the official repository from what is still only an announcement: the cumulative memory actually available, measured speed gains on Apple hardware, network cost between machines, and the precise matrix of supported platforms before investing in multiple devices.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What Exo actually does

Exo (the exo-explore/exo repository, with nearly 48,000 GitHub stars as of September 28, 2026) is described by its creators as a tool to “run cutting-edge AI locally.” Devices running Exo discover each other automatically on the network, without manual configuration, and expose a dashboard and an API at http://localhost:52415 on each node.

The project highlights “topology-aware” tensor parallelism: Exo evaluates network topology in real time (latency and bandwidth between every pair of machines) and each device's resources to decide how to distribute a model, rather than applying the same fixed split regardless of the available hardware.

i
This is not multi-GPU in a single machine
Exo explicitly targets networking multiple separate computers. To distribute a model across multiple GPUs installed in a single machine, llama.cpp's tensor-split is the appropriate tool, not Exo.

#Cumulative memory: the real contribution

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The most tangible benefit of an Exo cluster is the sum of the available memory. One configuration illustrated by the project itself combines 4 Mac Studio M3 Ultra systems with 512 GB each to load DeepSeek v3.1 in 8-bit and Kimi-K2-Thinking in 4-bit simultaneously—models that none of these machines could load alone, even with 512 GB.

In terms of compute, Exo claims a gain from tensor parallelism: up to 1.8x faster on 2 devices and 3.2x on 4 devices. These are speedup factors for distributed computation, not tokens-per-second throughput: the actual number of tokens generated per second always depends on the model, quantization, and—as we’ll discuss below—the network between the machines.

#Network cost: what RDMA changes

Connecting machines over a network introduces latency that a single GPU doesn't have. Exo addresses this with RDMA support (remote direct memory access) over Thunderbolt 5, announced when the feature was released, claiming a 99% reduction in latency between devices compared with a standard network connection.

This RDMA capability is not universal: it relies on a feature added in macOS 26.2 and works only on Macs equipped with Thunderbolt 5 — Mac mini M4 Pro, Mac Studio M4 Max or M3 Ultra, MacBook Pro M4 Max according to the official list. The project also requires all devices in the RDMA cluster to be interconnected, with TB5-certified cables, and the macOS version (including beta versions) to be strictly identical on every machine; otherwise, the RDMA ports may fail to discover one another.

!
Without RDMA, the network remains the limiting factor
Outside the RDMA setup over Thunderbolt 5, an Exo cluster communicates over the standard network (Wi-Fi or Ethernet), whose latency and bandwidth directly affect generation throughput, especially for model layers that must exchange activations between machines at each step.

#Matrix of platforms that are actually supported

What Exo actually supports by platform, as of September 28, 2026
PlatformAccelerationStatus
macOS (Apple Silicon)GPU via MLXPrimary path, Thunderbolt 5 RDMA available on recent Macs running macOS 26.2+
Linux (x86/ARM)CPU onlyGPU support is under development according to the official README
WindowsNo official mentionNot documented in the repository as of September 28, 2026

This table contradicts a common assumption: Exo is not a tool that runs a heterogeneous Mac plus PC cluster at full speed. The MLX inference backend is specific to Apple Silicon; on Linux, the official README clearly states that Exo currently runs on the CPU, without GPU acceleration, with GPU support announced as “in development” and no delivery date specified.

!
Do not promise a high-performing mixed cluster
A cluster composed exclusively of Mac Apple Silicon remains the reference configuration documented by the project. Adding a Linux machine to the cluster is possible, but that machine will contribute CPU, not GPU, with an impact on overall speed that must be measured rather than assumed.

#Deploying a cluster in practice

  1. 01
    Prepare each macOS machine
    Install Xcode (for the Metal toolchain), Homebrew, uv, and Node—the documented prerequisites for building and running Exo from source on macOS.
  2. 02
    Clone and run Exo
    git clone du dépôt, construction du tableau de bord (npm install && npm run build), puis uv sync --extra mlx et uv run exo sur chaque machine.
  3. 03
    Check automatic discovery
    Devices on the same network discover one another without configuration; the dashboard available at http://localhost:52415 should list every active cluster node.
  4. 04
    Enable RDMA if the hardware supports it
    On Thunderbolt 5 Macs running macOS 26.2+, enable rdma_ctl enable from Recovery mode on each machine, then connect all the devices to one another with TB5-certified cables.
  5. 05
    Load a sharded model
    From the dashboard or API, choose a model whose size exceeds the memory of a single device to verify that partitioning works in practice before targeting even larger models.

#Compatible APIs and existing clients

Exo exposes several compatible APIs: OpenAI Chat Completions, OpenAI Responses, Claude Messages, and Ollama—a client already written for one of these formats can point to the Exo cluster without rewriting. It’s a pragmatic choice that avoids having to adopt a proprietary protocol to use the cluster.

Environment variables supplement the configuration for advanced use: EXO_OFFLINE to run without an internet connection once the models are already cached, EXO_MODELS_READ_ONLY_DIRS to share read-only model storage between machines (useful for a shared NFS mount), and EXO_LIBP2P_NAMESPACE to isolate multiple Exo clusters on the same physical network.

#Custom models from Hugging Face

Exo is not limited to a closed catalog of models: the project supports custom models directly from the Hugging Face Hub, beyond the demonstration deployments highlighted by its creators (DeepSeek v3.1 with 671 billion parameters, Qwen3-235B, Kimi-K2-Thinking). This makes it possible to test a recent model as soon as it is published on the Hub, without waiting for it to be added to an officially supported list.

This openness has a documented tradeoff: custom models that require trust_remote_code in their configuration must be explicitly enabled, as this setting remains disabled by default for security reasons. Running arbitrary code supplied by a third-party Hugging Face repository across all machines in the cluster is therefore not the default behavior, but a choice to make consciously for models that require it.

#Security: a cluster with no documented authentication

The official README and dashboard documentation mention no authentication mechanism, API key, token, or TLS encryption for the Exo API and dashboard: as documented on September 28, 2026, anyone who can reach port 52415 on a node can query the API or view the dashboard without a login step.

!
Trusted network, no direct exposure
Without documented authentication, an Exo cluster should remain on a trusted local network (home or internal network) rather than being exposed directly to the internet or to a shared network with unidentified third parties. Adding your own reverse proxy with authentication is a precaution for the user to take, not a feature provided by the project.

This point deserves even more attention because each machine in the cluster exposes its own API on the local network: in a multi-device setup, the exposed surface is that of each node individually, not just that of a single entry point that could be protected in isolation.

#Troubleshooting: symptoms, cause, fix

Common symptoms when setting up an Exo cluster
SymptomLikely causeCorrection
RDMA ports are not discoverable between two MacsDifferent macOS versions across machines, including beta versionsStrictly align the macOS version (including beta versions) across every device in the cluster
A custom model refuses to load from Hugging Facetrust_remote_code required by the model but disabled by defaultExplicitly enable trust_remote_code for this specific model, understanding the code it executes
Disappointing throughput despite adding several machinesNo RDMA: the cluster runs over Wi-Fi or standard Ethernet, making it more latency-sensitiveVerify RDMA eligibility (Thunderbolt 5, macOS 26.2+) or move the machines closer together on a dedicated wired network
A Linux machine in the cluster does not appear to accelerate inferenceGPU support on Linux is still under development; the machine contributes CPU onlyCount this machine only as memory and CPU compute, not as a GPU accelerator

#Limits and pitfalls to anticipate

No documented Windows support
The official repository mentions no Windows support as of September 28, 2026; the installation paths cover macOS and Linux only.
Linux limited to the CPU
GPU-accelerated inference on Linux is in development, with no announced release date; a Linux-only cluster will be significantly slower than an equivalent Apple Silicon cluster.
Hardware-intensive RDMA
Thunderbolt 5, macOS 26.2 or later, and strictly identical system versions on every machine: a single device that fails any of these requirements falls back to a slower conventional network.
Throughput not guaranteed
The acceleration factors (1.8x, 3.2x) measure the benefit of tensor parallelism for distributed computation, not a tokens-per-second throughput that can be directly transferred to your own model and network.
Frequently asked questions
Can Exo Mix Macs and Windows PCs in the Same Cluster?+
The official repository documents no Windows support. Only macOS and Linux have documented installation paths, so a mixed Mac and Windows PC cluster is not a supported configuration at this time, contrary to an assumption sometimes made about the project.
Does Exo use the GPU on Linux?+
No, not yet. The official README explicitly states that Exo currently runs CPU-only on Linux, with GPU support announced as in development and no delivery date specified. A Linux machine added to the cluster therefore contributes memory and CPU compute, not GPU acceleration.
Do you need RDMA over Thunderbolt to use Exo?+
No, Exo works over a standard network, Wi-Fi or Ethernet. RDMA over Thunderbolt 5, limited to certain recent Macs running macOS 26.2 or later, greatly reduces latency between machines but is not required to run a cluster—only to get the best possible throughput.
How much memory can an Exo cluster pool?+
It depends solely on the machines added: the project illustrates a configuration with 4 Mac Studio M3 Ultra systems, each with 512 GB, to load models such as DeepSeek v3.1 or Kimi-K2-Thinking, but the total combined memory is simply the sum of the memory in each device connected to the cluster.
Are the speed gains advertised by Exo (1.8x, 3.2x) guaranteed on my hardware?+
No, these are factors measured by the project on its own test hardware for tensor parallelism, not a tokens-per-second throughput figure that can be transferred as-is. Actual throughput depends on the model, quantization, number of machines, and above all the quality of the network connection between them.
Are Exo’s dashboard and API protected by a password?+
No authentication, API key, or TLS encryption is documented in the official README for the API or dashboard. Anyone who reaches port 52415 on a node can therefore query it without logging in: an Exo cluster should remain on a trusted local network rather than be exposed directly to the internet.
Can Exo load any model published on Hugging Face?+
Exo supports custom models from the Hugging Face Hub beyond its demo list. However, models that require trust_remote_code in their configuration must be enabled explicitly; this setting remains disabled by default to prevent arbitrary code from running without confirmation.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.