Multi-GPU LLM with llama.cpp: tensor-split 2× RTX 3090
A RTX 4090 24 GB isn’t enough to run the best 2026 models in full-quality Q8 with a 256k-token context. Two used 3090s (48 GB) do the job for less. But you need to know how to split the model across the cards, choose between split layer and tensor parallel, and configure the hardware to keep it running reliably. This guide covers the three common tools (llama.cpp, Ollama, vLLM) and the pitfalls.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Why multi-GPU
- Models > single-card VRAM
- Qwen3-Coder 30B-A3B in Q8 = 32 GB, and with a 256k context, the KV cache pushes the total well beyond that. No single consumer GPU can handle it, except a RTX 5090 (32 GB)—and even then, with no room for a large context.
- Throughput for teams
- Two cards in tensor parallel via vLLM can serve 2× more users concurrently.
- Redundancy
- One card fails, and the other continues in degraded mode (smaller model).
- Fine-tuning
- A LoRA on a 24B requires 30 GB for the model + optimizer + gradients. Two 24 GB cards = comfortable.
#Two strategies: split layer vs tensor parallel
- Layer split (layer parallelism)
- Layers 1–40 on GPU 0, 41–80 on GPU 1. Moving from one layer to the next copies an activation vector. Slightly higher latency. Supported by llama.cpp and Ollama.
- Tensor parallelism
- Each layer is split into chunks distributed across the GPUs in parallel. Requires fast communication between cards. More performant but complex. Supported by vLLM and ExLlamaV2.
#Hardware: watch the details
- PCIe
- Each GPU must have at least PCIe x8. Use a motherboard with 2× PCIe 5.0 x16 slots dedicated to the CPU, not x16 + x4 through the chipset.
- Power supply
- 2× 3090 = 700 W for the GPUs alone. Add the CPU, RAM, and SSD: target 1200–1600 W (80+ Gold or Platinum).
- Spacing
- The RTX 3090/4090 models are 3-slot. On a 7-slot motherboard, the two cards touch. Use PCIe risers or water cooling to keep access.
- Cooling
- With 2 cards under load, a poorly ventilated case reaches 85–90°C. Thermal throttling = −30% performance.
#1. Multi-GPU with llama.cpp
--tensor-split specifies the proportion of layers assigned to each GPU. For 2 identical cards: 0.5,0.5. For a 3090 24 GB + a 4070 12 GB: 0.67,0.33 (the 3090 takes two-thirds).
#2. Multi-GPU with Ollama
Ollama automatically detects multiple GPUs and distributes work across all available ones. Few exposed settings, but it works out of the box.
#3. Tensor parallelism with vLLM
vLLM distributes each layer across the 2 GPUs in parallel. Aggregate throughput is nearly doubled compared with 1 GPU, with about 10% communication overhead.
#NVLink: useful or not?
NVLink provides a high-bandwidth GPU-to-GPU connection (112 GB/s on the 3090). Some RTX 3090 have an NVLink connector; the 4090 no longer does; all RTX Quadro / A-series cards have one.
- Impact of llama.cpp split-layer
- Very low. Transfers between layers are small.
- Impact of vLLM tensor parallelism
- More significantly: +5 to +15% throughput with NVLink enabled.
- Fine-tuning impact
- Big picture: gradient communication is intensive. NVLink is recommended if you plan to fine-tune across multiple GPUs.
#Expected performance
On 2× RTX 3090, Qwen3-Coder 30B-A3B Q8_0, llama.cpp with layer splitting:
- Single prompt
- ~55 tok/s. A single 3090 couldn't fit the model in Q8 (32 GB). Speed remains high: it's a MoE with 3B active parameters.
- Time to first token
- ~200 ms (vs ~80 ms on a small single-GPU model).
- Same cards with vLLM tensor parallelism
- ~65 tok/s single, but ~400 tok/s aggregated across 10 parallel requests.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.