Advanced 22 minPerformance

Multi-GPU LLM with llama.cpp: tensor-split 2× RTX 3090

A RTX 4090 24 GB isn’t enough to run the best 2026 models in full-quality Q8 with a 256k-token context. Two used 3090s (48 GB) do the job for less. But you need to know how to split the model across the cards, choose between split layer and tensor parallel, and configure the hardware to keep it running reliably. This guide covers the three common tools (llama.cpp, Ollama, vLLM) and the pitfalls.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux
Recommended hardware

Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).

A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.

Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Compare all options by budget, from €800 to €3,500 →

Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.

On the go: which laptop for local AI →

Affiliate links — possible commission at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

#Why multi-GPU

Models > single-card VRAM
Qwen3-Coder 30B-A3B in Q8 = 32 GB, and with a 256k context, the KV cache pushes the total well beyond that. No single consumer GPU can handle it, except a RTX 5090 (32 GB)—and even then, with no room for a large context.
Throughput for teams
Two cards in tensor parallel via vLLM can serve 2× more users concurrently.
Redundancy
One card fails, and the other continues in degraded mode (smaller model).
Fine-tuning
A LoRA on a 24B requires 30 GB for the model + optimizer + gradients. Two 24 GB cards = comfortable.

#Two strategies: split layer vs tensor parallel

Layer split (layer parallelism)
Layers 1–40 on GPU 0, 41–80 on GPU 1. Moving from one layer to the next copies an activation vector. Slightly higher latency. Supported by llama.cpp and Ollama.
Tensor parallelism
Each layer is split into chunks distributed across the GPUs in parallel. Requires fast communication between cards. More performant but complex. Supported by vLLM and ExLlamaV2.
i
Rule of thumb
For single-user use and simplicity: split by layer (llama.cpp/Ollama). For multi-user serving and throughput: tensor parallelism (vLLM).

#Hardware: watch the details

PCIe
Each GPU must have at least PCIe x8. Use a motherboard with 2× PCIe 5.0 x16 slots dedicated to the CPU, not x16 + x4 through the chipset.
Power supply
2× 3090 = 700 W for the GPUs alone. Add the CPU, RAM, and SSD: target 1200–1600 W (80+ Gold or Platinum).
Spacing
The RTX 3090/4090 models are 3-slot. On a 7-slot motherboard, the two cards touch. Use PCIe risers or water cooling to keep access.
Cooling
With 2 cards under load, a poorly ventilated case reaches 85–90°C. Thermal throttling = −30% performance.

#1. Multi-GPU with llama.cpp

Split layer
./build/bin/llama-cli \
  -m ./models/qwen3-coder-30b-a3b-q8_0.gguf \
  -ngl 99 \
  --split-mode layer \
  --tensor-split 0.5,0.5 \
  -p "..."

--tensor-split specifies the proportion of layers assigned to each GPU. For 2 identical cards: 0.5,0.5. For a 3090 24 GB + a 4070 12 GB: 0.67,0.33 (the 3090 takes two-thirds).

Row split (sometimes faster)
./build/bin/llama-cli \
  -m ./models/qwen3-coder-30b-a3b-q8_0.gguf \
  -ngl 99 \
  --split-mode row \
  ...

#2. Multi-GPU with Ollama

Ollama automatically detects multiple GPUs and distributes work across all available ones. Few exposed settings, but it works out of the box.

Useful variables
# Limiter à certains GPU
export CUDA_VISIBLE_DEVICES=0,1

# Ou pour AMD
export ROCR_VISIBLE_DEVICES=0,1

systemctl restart ollama
Check the split
ollama ps
# Doit afficher le modèle réparti : PROCESSOR = 100% GPU,
# réparti sur les deux cartes vu par nvidia-smi

#3. Tensor parallelism with vLLM

Launch vLLM with 2 GPUs
python -m vllm.entrypoints.openai.api_server \
  --model Qwen/Qwen3-Coder-30B-A3B-Instruct \
  --tensor-parallel-size 2 \
  --quantization awq \
  --dtype auto \
  --port 8000

vLLM distributes each layer across the 2 GPUs in parallel. Aggregate throughput is nearly doubled compared with 1 GPU, with about 10% communication overhead.

→
Identical cards for vLLM
Tensor parallel works best with identical cards. 2× RTX 3090 or 2× 4090: OK. A 3090 + a 4070: vLLM will use the slower card as the baseline, leaving the faster one underutilized.

NVLink provides a high-bandwidth GPU-to-GPU connection (112 GB/s on the 3090). Some RTX 3090 have an NVLink connector; the 4090 no longer does; all RTX Quadro / A-series cards have one.

Impact of llama.cpp split-layer
Very low. Transfers between layers are small.
Impact of vLLM tensor parallelism
More significantly: +5 to +15% throughput with NVLink enabled.
Fine-tuning impact
Big picture: gradient communication is intensive. NVLink is recommended if you plan to fine-tune across multiple GPUs.

#Expected performance

On 2× RTX 3090, Qwen3-Coder 30B-A3B Q8_0, llama.cpp with layer splitting:

Single prompt
~55 tok/s. A single 3090 couldn't fit the model in Q8 (32 GB). Speed remains high: it's a MoE with 3B active parameters.
Time to first token
~200 ms (vs ~80 ms on a small single-GPU model).
Same cards with vLLM tensor parallelism
~65 tok/s single, but ~400 tok/s aggregated across 10 parallel requests.
!
PCIe bottleneck
If one of your cards is running at PCIe x4 (for example, a secondary slot connected through the chipset), you may lose 40% performance with split layer and even more with tensor parallelism. Check with nvidia-smi or GPU-Z before blaming everything else.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.