LLM on NVIDIA Jetson Orin: embedded AI that fits in the main
The NVIDIA Jetson Orin is the board that puts a real CUDA GPU into a case the size of a deck of cards. Unlike a Raspberry Pi, the Jetson Orin runs an LLM with hardware acceleration, which changes everything for edge AI: robotics, home automation, and autonomous sensors. This guide covers choosing the board (Nano 8 GB or AGX 64 GB), installing JetPack, deploying Ollama and llama.cpp, and selecting realistic models for your available memory—without ever leaving the local environment.
Choosing a machine? Our picks by budget →
Good value for local AI: a GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395).
A mini PC is a complete machine: check the available memory and engine compatibility. It does not replace macOS/MLX or CUDA.
Why this choice? Our complete guide on GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →
Compare all options by budget, from €800 to €3,500 →
Small budget: RTX 5060 · Large models: RTX 5090 · Mac Studio.
On the go: which laptop for local AI →
Affiliate links — commission possible at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.
#Why a Jetson Orin for an LLM
A Jetson Orin is not a microcontroller: it’s a system-on-chip module that combines an ARM Cortex CPU, a NVIDIA GPU with CUDA and Tensor cores, and unified memory shared by both. In practice, the GPU accesses system RAM directly—there is no separate VRAM. A loaded LLM therefore occupies unified memory, exactly as on an Apple Silicon Mac, and computation benefits from CUDA acceleration.
That’s what radically sets the Jetson Orin LLM apart from AI on a Raspberry Pi: where the Pi struggles on pure CPU, the Orin offloads to the GPU and delivers usable throughput on 3B to 8B models. All within a 7- to 60-watt power envelope, battery-powered, with no noisy fan or tower under your desk.
- Network independence
- No cloud dependency: the model runs onboard, including on a mobile robot or an isolated site without a connection.
- Predictable latency
- No internet round trip. The response depends only on the local GPU, which matters for a real-time control loop.
- Privacy
- Sensor, camera, and microphone data never leave the device. Essential for home automation and medical environments.
- CUDA ecosystem
- The same software stack as a NVIDIA desktop GPU: Ollama, llama.cpp, PyTorch, and TensorRT work, apart from ARM compilation.
#Jetson Orin Nano or AGX: choosing your board
The Orin lineup spans an 8x memory factor and far more in compute power. For an LLM, the decisive criterion is unified RAM: it caps the model size exactly as VRAM does on a conventional GPU. Here are the tiers that matter.
- Orin Nano 8 GB
- The entry point (~40 TOPS). 8 GB of unified RAM shared with the system. Targets 3B models in Q4, or a tight 7B if you free up memory. Ideal for a smart sensor or embedded voice assistant.
- Orin NX 8 / 16 GB
- Mid-range (~70-100 TOPS). The 16 GB comfortably handles 7B-8B models in Q4_K_M. A good compromise for mobile robotics.
- AGX Orin 32 GB
- Embedded workstation (~200 TOPS). Runs a 14B in Q4 with context, or several small models in parallel (vision + language).
- AGX Orin 64 GB
- The high-end option (~275 TOPS). Targets 32B models in Q4 (~19 GB), with room for context and other workloads. The only one in the lineup suited to a serious 32B.
#JetPack prerequisites and role
The entire Jetson ecosystem is built on JetPack, NVIDIA's software distribution. It bundles Ubuntu (L4T, Linux for Tegra), CUDA drivers, cuDNN, TensorRT, and GPU libraries. Without JetPack flashed correctly, there is no acceleration: you'd fall back to pure CPU, losing the point of the card.
- A Jetson Orin card
- Nano, NX, or AGX, with the appropriate power supply (the Nano draws up to 15 W, and the AGX up to 60 W).
- Fast storage
- A microSD card (UHS-I) for the Nano during testing, but an NVMe SSD is strongly recommended: the models weigh several GB, and loading suffers on SD.
- A Linux host PC (Ubuntu)
- Required for flashing via SDK Manager on AGX/NX. The Nano Developer Kit installs directly from an SD image.
- JetPack 6.x
- The recent branch is based on Ubuntu 22.04 with CUDA 12. This is the target for a modern LLM stack (Ollama and llama.cpp compiled for the ARM SBSA architecture).
#1. Flash JetPack and verify CUDA
- 01Flash the imageFor an Orin Nano Developer Kit, write the official SD image to a card with Balena Etcher, insert it, connect the device, and follow the Ubuntu wizard. For an AGX/NX, use NVIDIA SDK Manager on the host PC, with a USB-C cable and the board in recovery mode.
- 02Update the systemOn the first startup, update the packages. This is also the time to install any missing JetPack components if you used the SD image.
- 03Check the GPU and CUDACheck that CUDA is present and that the module is recognized. jtop (installed via the jetson-stats package) provides a real-time view of the GPU, RAM, and power consumption—the Jetson equivalent of nvidia-smi.
#2. Install Ollama under JetPack
Ollama is the fastest route to your first model. The official installation script detects the ARM64 architecture and Jetson GPU, and configures the systemd service. Once started, the daemon listens on http://localhost:11434 by default, just as it does on a desktop PC.
For an interface like ChatGPT, add Open WebUI as a container pointing to the same endpoint. On an 8 GB Nano, keep it lean: every service consumes unified memory already counted against the model.
#3. Compile llama.cpp with CUDA
Ollama is enough for most use cases, but compiling llama.cpp manually gives you fine-grained control: choosing the exact GGUF quantization, adjusting the number of layers offloaded to the GPU, and often gaining a few more tokens/sec. It is also the way to go if you want to integrate the engine into your own embedded binary.
#Which models run with your memory
The rule is the same as on desktop GPUs: with Q4_K_M quantization (the best quality-to-memory tradeoff), plan for ~2 GB for a 3B, ~5 GB for a 7B, ~9 GB for a 14B, and ~19 GB for a 32B. On Jetson, subtract system memory from total RAM to determine your actual budget.
- Orin Nano 8 GB
- Granite 4.2 3B (~2.2 GB), Qwen 3.5 4B (~3.4 GB), or Gemma 4 E2B (~4.3 GB) in Q4. An 8B (Granite 4.2 8B, ~5.3 GB) works if you close everything else, but the context remains tight.
- Orin NX 16 GB
- Comfortable with 8B–9B models (Granite 4.2 8B, Qwen 3.5 9B) in Q4_K_M with some context. A 24B model (Mistral Small 24B, ~14 GB) fits in Q4 with little headroom.
- AGX Orin 32 GB
- A 24B (Mistral Small, gpt-oss 20B) comfortably, or a 30-35B MoE (Qwen 3.6 35B-A3B, ~23 GB). Enough room to combine a language model with a vision stack.
- AGX Orin 64 GB
- A full-quality 30–35B MoE, a 24B in Q8, or several models loaded simultaneously. The only tier that seriously targets large models with room to spare.
#Power consumption and power modes
The Jetson's key advantage is performance per watt. Each board offers power modes that cap power consumption by enabling more or fewer CPU cores and limiting GPU frequency. They're controlled with nvpmodel, and jetson_clocks forces the frequencies to the maximum allowed by the mode.
- Low-power mode
- On Orin Nano, a 7 W mode severely limits throughput but allows operation from a battery or modest USB-C power source. Suitable for a sensor that queries the model intermittently.
- Full-power mode
- MAXN mode unlocks the entire GPU. On AGX Orin, power consumption can reach 60 W: plan for cooling (an active heatsink) and a properly sized power supply.
- The embedded trade-off
- Many projects run well in 15-25 W mode: good throughput on a 3B-7B model, manageable heat dissipation, and reasonable battery life.
#Use case: robotics and home automation
The Jetson shines where an LLM needs to operate as close as possible to the physical world, without cloud latency or data leaks. Two areas dominate embedded applications.
- Robotics (ROS 2)
- The LLM serves as a natural-language interface: translating a spoken instruction into a sequence of actions, describing a scene perceived by the camera, and reasoning about a task. The Jetson hosts both perception (vision) and language on the same GPU.
- Local home automation
- A homegrown voice assistant that controls Home Assistant without ever using the cloud. The model interprets natural-language requests and triggers automations—your microphone and data stay at home.
- Autonomous smart sensor
- At an isolated site (agricultural, industrial), Orin analyzes data locally and generates natural-language summaries, transmitted only when a connection is available.
- Offline field assistant
- Technical documentation queryable with RAG, embedded in a vehicle or piece of equipment, operating without a network connection.
#Against the Raspberry Pi: the real gap
The question always comes up: why pay for a Jetson when a Raspberry Pi 5 can also run Ollama? The answer comes down to one word: GPU. The Pi has no usable accelerator for LLM inference; everything runs on the ARM CPU. The Jetson, meanwhile, offloads computation to CUDA cores.
- Acceleration
- Pi 5: CPU-only, managing just a few tokens/sec on a 3B. Jetson Orin: CUDA GPU, several times the throughput on the same model, with genuinely usable 7B-14B models.
- Memory
- The Pi 5 tops out at 8/16 GB of CPU RAM. Jetson supports up to 64 GB of unified memory usable by the GPU, enabling models beyond the Pi’s reach.
- Pricing
- The Pi costs a fraction of the Jetson. The price difference is real: the Jetson is justified when you need throughput or large models, not for a simple 1B bot.
- Ecosystem
- The Pi is general-purpose and community-driven; the Jetson targets embedded AI with CUDA, TensorRT, and NVIDIA robotics support (Isaac).
#Troubleshooting
- Ollama remains on the CPU
- The GPU isn't being used if JetPack/CUDA isn't installed correctly, or if you flashed a generic ARM image without Tegra support. Check nvcc --version and GPU activity in jtop.
- Out of memory while loading
- The model exceeds the available unified memory after accounting for the system. Drop down one size (7B → 3B), switch to a more aggressive Q4, or close other services.
- Throughput collapses
- Thermal throttling. Check the temperature in jtop, improve cooling, or switch to a less power-hungry but stable nvpmodel mode.
- Very slow model loading
- Models stored on microSD. Move ~/.ollama (or your GGUF files) to an NVMe SSD for loads in seconds rather than minutes.
- llama.cpp ignores the GPU
- The -ngl option is missing, or the build lacks CUDA. Recompile with -DGGML_CUDA=ON and launch with -ngl 99.
#Go further
The Jetson shares the same software stack as a NVIDIA PC. These guides naturally extend your embedded setup:
- LLM on Raspberry Pi 5: embedded local AI
- A direct comparison using pure CPU performance, useful for choosing between Pi and Jetson based on your project and budget.
- Compile llama.cpp with CUDA
- For a deeper look at building from source and GPU settings that can be applied to the Jetson’s ARM architecture.
- Choose your quantization (Q4, Q5, Q8, FP16)
- To determine the right memory footprint for your Orin board's unified RAM.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.