Ollama on Proxmox: LXC, VM, and GPU passthrough
Proxmox VE is the homelab Swiss Army knife: a free hypervisor that is light on resources and can run lightweight containers and isolated virtual machines side by side on a single server. Running Ollama on Proxmox lets you share an expensive GPU among multiple services while keeping inference 100% on-premises. This guide covers both approaches—LXC with shared GPU access (simple) and a VM with full passthrough (clean)—with detailed NVIDIA configuration and the IOMMU pitfalls that can cost you entire evenings.
#Why host Ollama on Proxmox
In a homelab, you rarely want a server dedicated solely to AI. Proxmox lets you place Ollama alongside your other services (NAS, home automation, reverse proxy) on the same machine while isolating each workload. The Ollama daemon then runs in a container or VM, listens on its usual port, and becomes accessible from your entire local infrastructure—Open WebUI from another container, an n8n pipeline, or your workstation.
The central issue is the GPU. LLM inference without hardware acceleration is usable but slow; as soon as you want a 14B or 32B model at a reasonable speed, you need to give Ollama access to a NVIDIA card. But virtualizing a GPU is not trivial: Proxmox offers two radically different approaches, with opposing tradeoffs in terms of simplicity and cleanliness.
#LXC vs VM: which choice for Ollama
Agents that act on your machine: agentic Cline, MCP, n8n + Ollama, local automations.
- Lifetime online access
- PDF + files
- Lifetime updates
This is the first decision, and it determines everything else. An LXC container shares the host kernel: it is lightweight, starts instantly, and above all can access the host's GPU without taking it away. A VM, by contrast, is a complete and isolated machine; to give it a GPU, you have to dedicate the entire GPU to it through VFIO passthrough—the host then loses that card completely.
- LXC + shared GPU
- The simple approach. The NVIDIA driver is installed on the host and in the container (same versions), and the /dev/nvidia* devices are mounted in the container. The GPU remains usable by the host and other containers simultaneously. Ideal if you want to share a single card.
- VM + full passthrough
- The clean approach. The GPU is detached from the host and assigned to the VM through VFIO/IOMMU. Complete isolation, drivers managed in the VM as with bare metal, and no interference. But the GPU is monopolized by that VM for as long as it is running.
- The deciding factor
- One card shared between multiple services → LXC. A card dedicated to AI (or multiple GPUs) with strict isolation required → VM. Passthrough is also mandatory if you want to run a different OS (Windows, or an OS with specific drivers).
#Hardware and VRAM requirements
Size the GPU according to the models you plan to run. At Q4_K_M quantization (the best quality/memory tradeoff), here is the required VRAM for each model size. Always leave some headroom for the context.
- 3B (≈2 GB)
- RTX 3060 12GB is more than sufficient. Comfortable even in CPU-only mode for testing.
- 7B (≈5 GB)
- RTX 3060 12GB or RTX 4070 12GB. The ideal entry point for a homelab.
- 14B (≈9 GB)
- RTX 4070 12GB is tight, RTX 4080 16GB is comfortable.
- 32B (≈19 GB)
- RTX 4090 24GB required, or two cards in tensor split.
- 70B (≈40 GB)
- Multi-GPU (2× RTX 4090) or a Mac M4 Pro/Max with 48 GB+ of unified memory.
For virtualization, VM passthrough requires strict prerequisites that LXC does not: a CPU and motherboard supporting VT-d (Intel) or AMD-Vi (AMD), IOMMU enabled in the BIOS, and ideally a clean IOMMU group for the GPU. LXC needs none of that—only for the NVIDIA driver to be installed on the host.
#LXC with shared GPU, the simple route
The principle: install the NVIDIA driver on the Proxmox host, create an unprivileged LXC container, mount the GPU devices in it, then install the same driver (without the kernel module) and Ollama in the container. The driver version must be identical on the host and in the container; otherwise, the container's userspace refuses to communicate with the host's kernel module.
- 01Install driver NVIDIA on the hostAdd the kernel headers, then install the driver from the NVIDIA repository (or the official .run file). Verify with nvidia-smi that the card is detected correctly on the host before proceeding.
- 02Create an unprivileged LXC containerA Debian 12 or Ubuntu 24.04 template is enough. Enable nesting (nesting=1) to allow GPU runtimes to run. Note the container ID (e.g., 200).
- 03Identify GPU devicesOn the host, list the /dev/nvidia* nodes and note their major/minor numbers. These are the device nodes that will be exposed to the container.
- 04Mount devices in the LXC configurationEdit /etc/pve/lxc/200.conf to allow NVIDIA device cgroups and bind-mount them. Restart the container.
- 05Install the same driver in the containerIn the container, install the NVIDIA driver with the --no-kernel-module option (the module comes from the host). nvidia-smi should now work inside it.
- 06Install OllamaThe official script detects the GPU automatically. Ollama starts its daemon and loads the models onto the shared card.
#VM with full passthrough, the clean approach
Here, the GPU is detached from the host and assigned to a VM via VFIO. The VM sees a real NVIDIA card and behaves like bare metal: install the driver normally, with no version workaround. In return, you must prepare the host to release the card at startup (VFIO binding) and ensure the kernel does not load the nvidia driver on the host.
- 01Enable IOMMU at bootAdd intel_iommu=on (or amd_iommu=on) and iommu=pt to the kernel parameters, in GRUB or systemd-boot depending on your Proxmox installation.
- 02Isolate the GPU for VFIOIdentify the card’s PCI IDs (GPU + HDMI audio), declare them in vfio-pci, and blacklist the nouveau/nvidia drivers on the host so they do not claim the device.
- 03Regenerate the initramfs and rebootUpdate the initramfs to include the VFIO modules, then reboot the host. Verify that the GPU is properly handled by vfio-pci.
- 04Create the VM and assign it the GPUCreate a VM (q35 machine type, OVMF/UEFI BIOS), add the GPU's PCI device in passthrough mode, and check PCI Express. Install the guest OS (Ubuntu Server recommended).
- 05Install driver + Ollama in the VMIn the VM, install the standard NVIDIA driver (the kernel module is allowed here; this is a real machine), verify with nvidia-smi, then install Ollama.
#IOMMU, drivers, and common pitfalls
Passthrough rarely fails randomly: it’s almost always caused by the IOMMU or a poorly isolated PCI group. Here are the most common recurring blockers and how to diagnose them.
- Non-isolated IOMMU groups
- If the GPU shares its IOMMU group with other devices (USB controller, another card), Proxmox refuses passthrough. Check the groups; as a last resort, the ACS override patch separates the groups, at the cost of less strict isolation.
- The host takes over the GPU
- If, after rebooting, lspci shows the GPU under “nvidia” or “nouveau” rather than “vfio-pci,” the blacklist did not take effect. Check /etc/modprobe.d and regenerate the initramfs.
- IOMMU not enabled in the BIOS
- dmesg shows no DMAR/IOMMU lines: VT-d or AMD-Vi is disabled in the firmware. Enable it in the BIOS first.
- LXC driver version out of sync
- LXC side only: NVML mismatch after a host update. Align the host and container versions.
- GPU used by the host console
- A single GPU also serving as Proxmox's video output is difficult to pass through cleanly. Ideally, keep an iGPU or a second card for displaying the host.
#Expose Ollama on the local network
By default, Ollama listens only on http://localhost:11434—so only from inside the container or VM. To call it from Open WebUI hosted elsewhere, or from your workstation, you must tell it to listen on all interfaces through the OLLAMA_HOST variable.
#Troubleshooting
- nvidia-smi missing from the container
- Devices not mounted or cgroups not properly authorized. Check the lxc.mount.entry and lxc.cgroup2.devices.allow lines, then verify the actual major numbers with ls -l /dev/nvidia*.
- Ollama runs on the CPU despite the GPU
- Check ollama ps: if the model is “100% CPU,” the CUDA runtime cannot see the card. Check nvidia-smi and restart the Ollama daemon.
- VM won’t start after PCI device addition
- Often an OVMF/q35 issue or a shared IOMMU group. Check the VM logs, and verify that the card is using vfio-pci on the host.
- Degraded performance after a Proxmox update
- A host kernel update can break the NVIDIA (LXC) module or VFIO binding. Reinstall the headers and driver, then regenerate the initramfs.
#Go further
Once Ollama is set up on Proxmox, these site guides help you finalize the stack and choose the right hardware:
- Install Ollama on Linux
- Covers clean daemon installation, systemd, and NVIDIA/AMD GPU configuration in detail — useful inside the container or VM.
- Deploy an LLM in production with Docker Compose
- To fit Open WebUI, Qdrant, and a reverse proxy into your Ollama Proxmox in a complete stack.
- Choose your GPU for local AI
- The 2026 buying guide for sizing the card to deploy in LXC or passthrough, depending on the models you target.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.