Advanced 15 minProxmox

Ollama on Proxmox: LXC, VM, and GPU passthrough

Proxmox VE is the homelab Swiss Army knife: a free hypervisor that is light on resources and can run lightweight containers and isolated virtual machines side by side on a single server. Running Ollama on Proxmox lets you share an expensive GPU among multiple services while keeping inference 100% on-premises. This guide covers both approaches—LXC with shared GPU access (simple) and a VM with full passthrough (clean)—with detailed NVIDIA configuration and the IOMMU pitfalls that can cost you entire evenings.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why host Ollama on Proxmox

In a homelab, you rarely want a server dedicated solely to AI. Proxmox lets you place Ollama alongside your other services (NAS, home automation, reverse proxy) on the same machine while isolating each workload. The Ollama daemon then runs in a container or VM, listens on its usual port, and becomes accessible from your entire local infrastructure—Open WebUI from another container, an n8n pipeline, or your workstation.

The central issue is the GPU. LLM inference without hardware acceleration is usable but slow; as soon as you want a 14B or 32B model at a reasonable speed, you need to give Ollama access to a NVIDIA card. But virtualizing a GPU is not trivial: Proxmox offers two radically different approaches, with opposing tradeoffs in terms of simplicity and cleanliness.

i
What this guide assumes
Proxmox VE 8.x installed and working, a NVIDIA GPU (RTX 3060 12GB or better), and root access to the node via SSH or the web console. The commands are provided for a Debian 12 host (the basis of Proxmox 8) and a Debian/Ubuntu guest.

#LXC vs VM: which choice for Ollama

The Local Agents Kit

Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

This is the first decision, and it determines everything else. An LXC container shares the host kernel: it is lightweight, starts instantly, and above all can access the host's GPU without taking it away. A VM, by contrast, is a complete and isolated machine; to give it a GPU, you have to dedicate the entire GPU to it through VFIO passthrough—the host then loses that card completely.

LXC + shared GPU
The simple approach. The NVIDIA driver is installed on the host and in the container (same versions), and the /dev/nvidia* devices are mounted in the container. The GPU remains usable by the host and other containers simultaneously. Ideal if you want to share a single card.
VM + full passthrough
The clean approach. The GPU is detached from the host and assigned to the VM through VFIO/IOMMU. Complete isolation, drivers managed in the VM as with bare metal, and no interference. But the GPU is monopolized by that VM for as long as it is running.
The deciding factor
One card shared between multiple services → LXC. A card dedicated to AI (or multiple GPUs) with strict isolation required → VM. Passthrough is also mandatory if you want to run a different OS (Windows, or an OS with specific drivers).
→
The default recommendation
For a first setup with a single GPU in a homelab, start with the LXC. It's faster to set up, has fewer pitfalls, and keeps the card available for other uses. Move to a VM when you have a real need for isolation or a dedicated GPU.

#Hardware and VRAM requirements

Size the GPU according to the models you plan to run. At Q4_K_M quantization (the best quality/memory tradeoff), here is the required VRAM for each model size. Always leave some headroom for the context.

3B (≈2 GB)
RTX 3060 12GB is more than sufficient. Comfortable even in CPU-only mode for testing.
7B (≈5 GB)
RTX 3060 12GB or RTX 4070 12GB. The ideal entry point for a homelab.
14B (≈9 GB)
RTX 4070 12GB is tight, RTX 4080 16GB is comfortable.
32B (≈19 GB)
RTX 4090 24GB required, or two cards in tensor split.
70B (≈40 GB)
Multi-GPU (2× RTX 4090) or a Mac M4 Pro/Max with 48 GB+ of unified memory.

For virtualization, VM passthrough requires strict prerequisites that LXC does not: a CPU and motherboard supporting VT-d (Intel) or AMD-Vi (AMD), IOMMU enabled in the BIOS, and ideally a clean IOMMU group for the GPU. LXC needs none of that—only for the NVIDIA driver to be installed on the host.


#LXC with shared GPU, the simple route

The principle: install the NVIDIA driver on the Proxmox host, create an unprivileged LXC container, mount the GPU devices in it, then install the same driver (without the kernel module) and Ollama in the container. The driver version must be identical on the host and in the container; otherwise, the container's userspace refuses to communicate with the host's kernel module.

  1. 01
    Install driver NVIDIA on the host
    Add the kernel headers, then install the driver from the NVIDIA repository (or the official .run file). Verify with nvidia-smi that the card is detected correctly on the host before proceeding.
  2. 02
    Create an unprivileged LXC container
    A Debian 12 or Ubuntu 24.04 template is enough. Enable nesting (nesting=1) to allow GPU runtimes to run. Note the container ID (e.g., 200).
  3. 03
    Identify GPU devices
    On the host, list the /dev/nvidia* nodes and note their major/minor numbers. These are the device nodes that will be exposed to the container.
  4. 04
    Mount devices in the LXC configuration
    Edit /etc/pve/lxc/200.conf to allow NVIDIA device cgroups and bind-mount them. Restart the container.
  5. 05
    Install the same driver in the container
    In the container, install the NVIDIA driver with the --no-kernel-module option (the module comes from the host). nvidia-smi should now work inside it.
  6. 06
    Install Ollama
    The official script detects the GPU automatically. Ollama starts its daemon and loads the models onto the shared card.
Proxmox host
# 1. En-têtes noyau + driver NVIDIA sur l'hôte
apt update && apt install -y pve-headers-$(uname -r)
apt install -y nvidia-driver nvidia-smi

# Vérifier la détection
nvidia-smi

# 2. Repérer les device nodes (relever major:minor)
ls -l /dev/nvidia*
/etc/pve/lxc/200.conf
# Conteneur non privilégié + nesting
features: nesting=1

# Autoriser les cgroups des périphériques NVIDIA (major 195, 234, 508 selon setup)
lxc.cgroup2.devices.allow: c 195:* rwm
lxc.cgroup2.devices.allow: c 234:* rwm
lxc.cgroup2.devices.allow: c 509:* rwm

# Monter les device nodes dans le conteneur
lxc.mount.entry: /dev/nvidia0 dev/nvidia0 none bind,optional,create=file
lxc.mount.entry: /dev/nvidiactl dev/nvidiactl none bind,optional,create=file
lxc.mount.entry: /dev/nvidia-uvm dev/nvidia-uvm none bind,optional,create=file
lxc.mount.entry: /dev/nvidia-uvm-tools dev/nvidia-uvm-tools none bind,optional,create=file
In the LXC container
# Même version de driver, SANS le module noyau (fourni par l'hôte)
./NVIDIA-Linux-x86_64-<version>.run --no-kernel-module

# nvidia-smi doit répondre à l'intérieur du conteneur
nvidia-smi

# Installer Ollama (détecte le GPU tout seul)
curl -fsSL https://ollama.com/install.sh | sh

# Test : le modèle doit se charger sur le GPU
ollama run qwen3.5:9b
!
The driver-version trap
If nvidia-smi displays "Failed to initialize NVML: Driver/library version mismatch" in the container, the container driver does not exactly match the host driver. Reinstall the identical version, or after a host update, restart the container to resynchronize.

#VM with full passthrough, the clean approach

Here, the GPU is detached from the host and assigned to a VM via VFIO. The VM sees a real NVIDIA card and behaves like bare metal: install the driver normally, with no version workaround. In return, you must prepare the host to release the card at startup (VFIO binding) and ensure the kernel does not load the nvidia driver on the host.

  1. 01
    Enable IOMMU at boot
    Add intel_iommu=on (or amd_iommu=on) and iommu=pt to the kernel parameters, in GRUB or systemd-boot depending on your Proxmox installation.
  2. 02
    Isolate the GPU for VFIO
    Identify the card’s PCI IDs (GPU + HDMI audio), declare them in vfio-pci, and blacklist the nouveau/nvidia drivers on the host so they do not claim the device.
  3. 03
    Regenerate the initramfs and reboot
    Update the initramfs to include the VFIO modules, then reboot the host. Verify that the GPU is properly handled by vfio-pci.
  4. 04
    Create the VM and assign it the GPU
    Create a VM (q35 machine type, OVMF/UEFI BIOS), add the GPU's PCI device in passthrough mode, and check PCI Express. Install the guest OS (Ubuntu Server recommended).
  5. 05
    Install driver + Ollama in the VM
    In the VM, install the standard NVIDIA driver (the kernel module is allowed here; this is a real machine), verify with nvidia-smi, then install Ollama.
Host — enable IOMMU (GRUB)
# Éditer /etc/default/grub, ligne GRUB_CMDLINE_LINUX_DEFAULT
# Intel :
GRUB_CMDLINE_LINUX_DEFAULT="quiet intel_iommu=on iommu=pt"
# AMD :
# GRUB_CMDLINE_LINUX_DEFAULT="quiet amd_iommu=on iommu=pt"

update-grub
reboot

# Après reboot : vérifier que l'IOMMU est actif
dmesg | grep -e DMAR -e IOMMU
Host — VFIO binding
# Trouver les IDs PCI et les vendor:device IDs du GPU
lspci -nn | grep -i nvidia
# ex. 01:00.0 ... [10de:2504]  (GPU)
#     01:00.1 ... [10de:228e]  (audio HDMI de la carte)

# Déclarer les IDs pour vfio-pci
echo "options vfio-pci ids=10de:2504,10de:228e" > /etc/modprobe.d/vfio.conf

# Blacklister les drivers côté hôte
echo -e "blacklist nouveau\nblacklist nvidia\nblacklist nvidiafb" > /etc/modprobe.d/blacklist-nvidia.conf

# Charger vfio au boot puis régénérer l'initramfs
echo -e "vfio\nvfio_iommu_type1\nvfio_pci" >> /etc/modules
update-initramfs -u -k all
reboot

# Après reboot : le GPU doit utiliser vfio-pci
lspci -nnk -d 10de:2504
→
PCI added to the Proxmox interface
In the VM → Hardware → Add → PCI Device, choose the GPU. Check “All Functions” to include HDMI audio, and “PCI-Express” (requires a VM with machine type q35 and OVMF BIOS). With recent NVIDIA drivers, the famous “Code 43” workaround is no longer needed under a Linux guest.

#IOMMU, drivers, and common pitfalls

Passthrough rarely fails randomly: it’s almost always caused by the IOMMU or a poorly isolated PCI group. Here are the most common recurring blockers and how to diagnose them.

Non-isolated IOMMU groups
If the GPU shares its IOMMU group with other devices (USB controller, another card), Proxmox refuses passthrough. Check the groups; as a last resort, the ACS override patch separates the groups, at the cost of less strict isolation.
The host takes over the GPU
If, after rebooting, lspci shows the GPU under “nvidia” or “nouveau” rather than “vfio-pci,” the blacklist did not take effect. Check /etc/modprobe.d and regenerate the initramfs.
IOMMU not enabled in the BIOS
dmesg shows no DMAR/IOMMU lines: VT-d or AMD-Vi is disabled in the firmware. Enable it in the BIOS first.
LXC driver version out of sync
LXC side only: NVML mismatch after a host update. Align the host and container versions.
GPU used by the host console
A single GPU also serving as Proxmox's video output is difficult to pass through cleanly. Ideally, keep an iGPU or a second card for displaying the host.
Host — check the IOMMU groups
# Lister les groupes IOMMU et leurs périphériques
for g in /sys/kernel/iommu_groups/*; do
  echo "Groupe ${g##*/}:"
  for d in $g/devices/*; do
    echo -n "  "; lspci -nns "${d##*/}"
  done
done

# Le GPU (et son audio) doit idéalement être seul dans son groupe

#Expose Ollama on the local network

By default, Ollama listens only on http://localhost:11434—so only from inside the container or VM. To call it from Open WebUI hosted elsewhere, or from your workstation, you must tell it to listen on all interfaces through the OLLAMA_HOST variable.

In the container / VM
# Écouter sur toutes les interfaces (override du service systemd)
mkdir -p /etc/systemd/system/ollama.service.d
cat > /etc/systemd/system/ollama.service.d/override.conf <<'EOF'
[Service]
Environment="OLLAMA_HOST=0.0.0.0:11434"
EOF

systemctl daemon-reload
systemctl restart ollama

# Depuis un autre hôte du LAN, tester l'API
curl http://<ip-conteneur>:11434/api/tags
!
Do not expose Ollama bare on the internet
The Ollama API has no authentication. When listening on 0.0.0.0, keep it strictly on your local network or behind a reverse proxy with authentication (Traefik, Caddy). Never forward port 11434 directly from your router.

#Troubleshooting

nvidia-smi missing from the container
Devices not mounted or cgroups not properly authorized. Check the lxc.mount.entry and lxc.cgroup2.devices.allow lines, then verify the actual major numbers with ls -l /dev/nvidia*.
Ollama runs on the CPU despite the GPU
Check ollama ps: if the model is “100% CPU,” the CUDA runtime cannot see the card. Check nvidia-smi and restart the Ollama daemon.
VM won’t start after PCI device addition
Often an OVMF/q35 issue or a shared IOMMU group. Check the VM logs, and verify that the card is using vfio-pci on the host.
Degraded performance after a Proxmox update
A host kernel update can break the NVIDIA (LXC) module or VFIO binding. Reinstall the headers and driver, then regenerate the initramfs.
Quick checks
# Le GPU est-il vu par Ollama ?
ollama ps          # doit montrer un % GPU, pas 100% CPU

# État de la carte et VRAM utilisée
nvidia-smi

# Le daemon répond-il ?
curl http://localhost:11434/api/tags

#Go further

Once Ollama is set up on Proxmox, these site guides help you finalize the stack and choose the right hardware:

Install Ollama on Linux
Covers clean daemon installation, systemd, and NVIDIA/AMD GPU configuration in detail — useful inside the container or VM.
Deploy an LLM in production with Docker Compose
To fit Open WebUI, Qdrant, and a reverse proxy into your Ollama Proxmox in a complete stack.
Choose your GPU for local AI
The 2026 buying guide for sizing the card to deploy in LXC or passthrough, depending on the models you target.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.